Work / gls:work:e38b054f-b926-4410-b84d-4fd3249fd4f4
GQ-002 · 10× the Agent Layer
A standing research question about making a provider-neutral multi-agent operating layer materially more effective, efficient, observable, and self-improving.
Research contract · GQ-002
What changes would make the complete provider-neutral SISO agent layer ten times more effective per unit of time, attention, and compute?
SISO Open Source Foundation
active
partial
Review due · assessed 2026-08-02 · due 2026-09-01
Run one paired agent task with and without the evidence-aware context and receipt path, including retrieval and repair overhead in the comparison.
gls:release:b95eeca5-ed9d-48ac-858c-f6d62065cd29 · program-metadata-1.0.0+ea7f0e4
Not released — the selected Release contains question-program metadata, not an answer artifact.
hybrid
Where to spend the next unit of agent-layer engineering effort: fixing an existing path's economics versus substituting a model or harness.
Every recommendation in the leverage map carries a measurement taken on this fleet, not an estimate or a vendor claim. · Each recommendation states a numeric success threshold that a later measurement could fail to meet. · The map ranks by measured effect on tokens, latency, or attention per unit of delivered work — not by expected effort to implement. · At least one recommendation has been applied and re-measured, with the before and after both recorded.
The highest-ranked lever, once applied and re-measured, produces less than a 2x improvement on the metric it was ranked by. Measured on the Bifrost request log, not asserted. · The ranking is unstable: re-deriving it from the same log a week later reorders the top three, meaning it described a moment rather than a structure. · Aggregate gains are dominated by a single provider or harness, so the 'agent layer' framing is wrong and the question should be about that component instead. · The measured bottleneck turns out to be human review or decision latency rather than tokens, latency, or compute, in which case no change to the agent layer reaches 10x.
No before/after pair exists for any applied lever: scripts/minimax-cache-route.sh is written and tested but NOT APPLIED, so the largest predicted gain is unverified. · Latency and attention are named in the question but only token economics have been measured. The ranking currently covers one of three axes. · Cached-read rates are measured on the gateway log only. Whether they reflect provider-side billing is unconfirmed.
A provider's cached-read rate on the gateway log moves by more than 20 percentage points, changing the economics the ranking was built on. · A new model or provider is added to the gateway, which can reorder the leverage map without any code changing. · Any recommendation in the map is applied, requiring a re-measurement to confirm or refute its predicted threshold. · Request volume on any provider falls below 100 in the trailing window, at which point per-request averages stop being trustworthy.
public agent harnesses and orchestration systems · released SISO Agent Stack components · privacy-safe operational measurements · memory, routing, budget, and handoff mechanisms
A ranked leverage map whose recommendations each carry measurements, external analogues, an experiment, and a falsifiable success threshold.
Re-answer after material stack changes or when new measurements overturn the current bottleneck model.
public metadata only
Read the God Questions infrastructure constitution →
Assumptions · 2
Evidence-aware context improves representative agent-task outcomes enough to justify its retrieval and authoring overhead. Scope: evidence-aware agent context · Falsifier: Paired representative tasks show no material correctness, intervention, reconstruction, or cost improvement after repair overhead is included. · Depends on: none · Evidence: EC-GQ002-OPERATING-PLAN · Review: 2026-08-02
Privacy-safe causal lineage and observation receipts make routing and capability decisions more reconstructable than untyped session summaries. Scope: agent outcome reconstruction · Falsifier: Independent reviewers cannot reproduce a routing or promotion decision more reliably from typed receipts than from the existing evidence path. · Depends on: none · Evidence: EC-GQ002-SEED-REVIEW · Review: 2026-08-02
Evidence connections · 2
The operating plan connects question demand, evidence, reversible experiments, measured outcomes, and returned learning without assigning truth to runtimes. Owner: The Great Library of SISO · Supports: QA-GQ002-CONTEXT-VALUE · Challenges: none · Observed: 2026-08-02
The public Work records that operational evidence and the existing answer remain unpublished rather than copying sensitive measurements. Owner: SISO Evidence Engines · Supports: QA-GQ002-LINEAGE-VALUE · Challenges: none · Observed: 2026-08-02
Action and learning lineage · 1
The public frame requests evidence about agent-layer leverage while private measurements and execution authority remain owner-held. Owner role: question steward · Status: proposed · Recorded: 2026-08-02T00:34:00+07:00 · Truth: not applicable · Predecessors: none · Assumptions: QA-GQ002-CONTEXT-VALUE, QA-GQ002-LINEAGE-VALUE
Research sources
Source & upstream links
Relationships
SISO FoundryFoundry supplies the outside-in harness, coordination, memory, and execution-pattern landscape.
SISO KnowledgeSISO Knowledge preserves public source and privacy-safe longitudinal evidence.
SISO Evidence EnginesEvidence Engines separate measured facts, hypotheses, counter-evidence, and improvement proposals.
Evidence & receipts
The existing campaign answer and receipts were read; operational and machine-specific evidence remains outside the public Library.source-review:frontier-question-intake:gq-002:2026-08-01
Provenance
- Registry source
- registry/works/frontier-question-gq-002.json
- Origin
- siso
- License / redistribution
- NONE (not_applicable)