Work / gls:work:e38b054f-b926-4410-b84d-4fd3249fd4f4

GQ-002 · 10× the Agent Layer

A standing research question about making a provider-neutral multi-agent operating layer materially more effective, efficient, observable, and self-improving.

Type
research_question
Maturity
active
Section
Research

Research contract · GQ-002

Question

What changes would make the complete provider-neutral SISO agent layer ten times more effective per unit of time, attention, and compute?

Steward

SISO Open Source Foundation

Lifecycle status

active

Research state

partial

Freshness

Review due · assessed 2026-08-02 · due 2026-09-01

Next useful work

Run one paired agent task with and without the evidence-aware context and receipt path, including retrieval and repair overhead in the comparison.

Selected release

gls:release:b95eeca5-ed9d-48ac-858c-f6d62065cd29 · program-metadata-1.0.0+ea7f0e4

Public answer release

Not released — the selected Release contains question-program metadata, not an answer artifact.

Evidence mode

hybrid

Decision to change

Where to spend the next unit of agent-layer engineering effort: fixing an existing path's economics versus substituting a model or harness.

Success criteria

Every recommendation in the leverage map carries a measurement taken on this fleet, not an estimate or a vendor claim. · Each recommendation states a numeric success threshold that a later measurement could fail to meet. · The map ranks by measured effect on tokens, latency, or attention per unit of delivered work — not by expected effort to implement. · At least one recommendation has been applied and re-measured, with the before and after both recorded.

Falsifiers

The highest-ranked lever, once applied and re-measured, produces less than a 2x improvement on the metric it was ranked by. Measured on the Bifrost request log, not asserted. · The ranking is unstable: re-deriving it from the same log a week later reorders the top three, meaning it described a moment rather than a structure. · Aggregate gains are dominated by a single provider or harness, so the 'agent layer' framing is wrong and the question should be about that component instead. · The measured bottleneck turns out to be human review or decision latency rather than tokens, latency, or compute, in which case no change to the agent layer reaches 10x.

Evidence gaps

No before/after pair exists for any applied lever: scripts/minimax-cache-route.sh is written and tested but NOT APPLIED, so the largest predicted gain is unverified. · Latency and attention are named in the question but only token economics have been measured. The ranking currently covers one of three axes. · Cached-read rates are measured on the gateway log only. Whether they reflect provider-side billing is unconfirmed.

Watch triggers

A provider's cached-read rate on the gateway log moves by more than 20 percentage points, changing the economics the ranking was built on. · A new model or provider is added to the gateway, which can reorder the leverage map without any code changing. · Any recommendation in the map is applied, requiring a re-measurement to confirm or refute its predicted threshold. · Request volume on any provider falls below 100 in the trailing window, at which point per-request averages stop being trustworthy.

Source scopes

public agent harnesses and orchestration systems · released SISO Agent Stack components · privacy-safe operational measurements · memory, routing, budget, and handoff mechanisms

Answer shape

A ranked leverage map whose recommendations each carry measurements, external analogues, an experiment, and a falsifiable success threshold.

Refresh policy

Re-answer after material stack changes or when new measurements overturn the current bottleneck model.

Publication boundary

public metadata only

Read the God Questions infrastructure constitution →

Assumptions · 2

QA-GQ002-CONTEXT-VALUE · QA · active · low

Evidence-aware context improves representative agent-task outcomes enough to justify its retrieval and authoring overhead. Scope: evidence-aware agent context · Falsifier: Paired representative tasks show no material correctness, intervention, reconstruction, or cost improvement after repair overhead is included. · Depends on: none · Evidence: EC-GQ002-OPERATING-PLAN · Review: 2026-08-02

QA-GQ002-LINEAGE-VALUE · QA · active · low

Privacy-safe causal lineage and observation receipts make routing and capability decisions more reconstructable than untyped session summaries. Scope: agent outcome reconstruction · Falsifier: Independent reviewers cannot reproduce a routing or promotion decision more reliably from typed receipts than from the existing evidence path. · Depends on: none · Evidence: EC-GQ002-SEED-REVIEW · Review: 2026-08-02

Evidence connections · 2

EC-GQ002-OPERATING-PLAN · documentation · public

The operating plan connects question demand, evidence, reversible experiments, measured outcomes, and returned learning without assigning truth to runtimes. Owner: The Great Library of SISO · Supports: QA-GQ002-CONTEXT-VALUE · Challenges: none · Observed: 2026-08-02

EC-GQ002-SEED-REVIEW · source_review · owner_held

The public Work records that operational evidence and the existing answer remain unpublished rather than copying sensitive measurements. Owner: SISO Evidence Engines · Supports: QA-GQ002-LINEAGE-VALUE · Challenges: none · Observed: 2026-08-02

Action and learning lineage · 1

AL-GQ002-DEMAND · epistemic demand · demand only

The public frame requests evidence about agent-layer leverage while private measurements and execution authority remain owner-held. Owner role: question steward · Status: proposed · Recorded: 2026-08-02T00:34:00+07:00 · Truth: not applicable · Predecessors: none · Assumptions: QA-GQ002-CONTEXT-VALUE, QA-GQ002-LINEAGE-VALUE

Research sources

Source & upstream links

Relationships

integrates_with

SISO FoundryFoundry supplies the outside-in harness, coordination, memory, and execution-pattern landscape.

depends_on

SISO KnowledgeSISO Knowledge preserves public source and privacy-safe longitudinal evidence.

coordinates_via

SISO Evidence EnginesEvidence Engines separate measured facts, hypotheses, counter-evidence, and improvement proposals.

Evidence & receipts

source_review · 2026-08-01

The existing campaign answer and receipts were read; operational and machine-specific evidence remains outside the public Library.source-review:frontier-question-intake:gq-002:2026-08-01

Provenance

Registry source
registry/works/frontier-question-gq-002.json
Origin
siso
License / redistribution
NONE (not_applicable)