Work / gls:work:3b4c4fa6-31f2-4cfc-9542-00f32e200621

GQ-008 · Model Routing Evidence

A standing research question about which available model should own each work shape after capability, reliability, latency, quota, and cost are considered together.

Type
research_question
Maturity
active
Section
Research

Research contract · GQ-008

Question

Which available model should own each recurring work shape after quality, reliability, tool use, latency, quota, and cost are evaluated together?

Steward

SISO Open Source Foundation

Lifecycle status

active

Research state

partial

Freshness

Not declared

Next useful work

Define the minimum public-safe question program before scheduling research.

Selected release

gls:release:a5e56fcf-e19d-4ead-8c7d-9660d41ed47e · registry-seed-2026-08-01

Public answer release

Not released — the selected Release has no public answer artifact.

Evidence mode

hybrid

Decision to change

Which model owns each recurring work shape, and whether an observed cost gap justifies changing the model or fixing the gateway path it was measured through.

Success criteria

Every cell in the routing matrix carries an evidence grade and the request volume it was derived from, so a reader can tell a measured cell from an assumed one. · Each recommendation names the work shape it applies to, not just the model — a matrix that ranks models without scenarios has not answered the question asked. · Cost and latency figures are derived from the gateway request log by a re-runnable query, never from vendor pricing pages or estimates. · The matrix distinguishes provider capability from gateway configuration, so a defect in the path is not recorded as a property of the model.

Falsifiers

A cell's recommendation, once routed to in production, is reversed on re-measurement — the model it named is not in fact best for that work shape. · Two cells derived from fewer than 100 requests each disagree with cells derived from thousands, indicating the matrix is reporting sampling noise as routing signal. · A measured cost gap between providers disappears once gateway configuration is corrected, meaning the matrix ranked a misconfiguration rather than a model. This has already partially fired: MiniMax reads 0 cached tokens through Bifrost and 2,944 of 3,033 through the local proxy. · Quality or tool-use reliability cannot be measured on this fleet at all, in which case the matrix covers only cost and latency and must say so rather than implying a complete answer.

Evidence gaps

Quality and tool-use reliability are named in the question but unmeasured. Only cost, cache behaviour, and volume have been derived, so the matrix currently answers a narrower question than it poses. · The MiniMax cache defect is isolated but NOT FIXED: scripts/minimax-cache-route.sh is written and tested, awaiting a topology decision. Every MiniMax cost figure describes a known-broken path. · Latency is not yet extracted from the gateway log, so the 'latency' axis of the question is unevidenced. · Cached-read counts come from the gateway log; whether they match provider-side billing is unconfirmed.

Watch triggers

Any provider's cached-read rate on the gateway log moves by more than 20 percentage points. · A model is added, removed, or re-versioned on the gateway. · Gateway topology changes — including applying scripts/minimax-cache-route.sh, which would invalidate every MiniMax cost cell. · Trailing-window volume for any provider drops below 100 requests, at which point its per-request averages stop being trustworthy.

Source scopes

current public benchmark suites · first-party provider documentation · privacy-safe real-task outcomes · quota, latency, and cost constraints

Answer shape

A versioned scenario-by-model routing matrix plus a machine-readable table, evidence grade per cell, and explicit re-evaluation triggers.

Refresh policy

Re-answer when a tracked model, benchmark, price, quota, or material real-world task result changes.

Publication boundary

public metadata only

Read the God Questions infrastructure constitution →

Research sources

Source & upstream links

Relationships

integrates_with

SISO FoundryFoundry discovers current benchmarks, model cards, evaluations, and counter-evidence.

depends_on

SISO KnowledgeSISO Knowledge preserves comparable benchmark and task evidence over time.

coordinates_via

SISO Evidence EnginesEvidence Engines separate benchmark observations, operational constraints, and scenario-level routing judgments.

Evidence & receipts

source_review · 2026-08-01

The existing benchmark campaign, routing table, and source coverage were read; only publication-safe metadata was promoted.source-review:frontier-question-intake:gq-008:2026-08-01

Provenance

Registry source
registry/works/frontier-question-gq-008.json
Origin
siso
License / redistribution
NONE (not_applicable)