PAIRED A/B LAB

Test one variable.

Use ChatGPT or an encrypted provider connection to compare request settings while the task and instructions remain identical.

NO-ACCOUNT RECORDED MEASUREMENT AUDIT

Review the measurement before connecting a model.

Inspect the public 2026-08-16 demonstration context and its published token metrics. This is an audit of the recorded numbers, not a replay: the original output text was intentionally not retained, so no output is invented and no winner is declared.

Run a live paired test
PUBLISHED TOKEN METRICS

The two recorded metric rows

GPT-5.5 · 2026-08-16. These are the public demonstration's reported numbers, shown separately from the unavailable output text.

Recorded arm A

Medium reasoning effort
Input
63
Output
109
Reasoning
20
Total
172

Recorded arm B

Low reasoning effort
Input
63
Output
94
Reasoning
0
Total
157
OUTPUT EVIDENCE BOUNDARY

Metrics are inspectable. Quality is not.

The public record says both outputs contained the requested three bullets, but the output text was intentionally not retained. A reviewer cannot independently judge their quality from this record, so TokenGauge does not declare a winner.

A valid live comparison: define a passing rubric first, judge both outputs before revealing usage, and count regressions or retries against any token reduction.

Inspect the current lab fixture and provenance
SHARED TASK · CONTEXT ONLY

Current starter wording

Explain why prompt-prefix stability matters to an engineering manager in three concise bullets.
SHARED INSTRUCTIONS · CONTEXT ONLY

Current starter wording

You are a helpful AI assistant. Preserve every required fact.

Current committed lab starter wording shown for context only; the exact 2026-08-16 demonstration prompt was not published.

Connect a model source to run a paired test.

3 controlled request-setting recipes are available. Use ChatGPT for the starter lab, or connect any supported provider API key from Settings with Pro.

Open provider settings