Measurement & integrity
Final rollback remains unconditional. Critical-clause shielding is now the default, with an explicit benchmark-only off ablation.
Promoted defaultsEvidence ledger · Updated July 15, 2026
A living record of what we tried, what actually saved tokens, what stayed intact, and what still needs proof. Experiments graduate only when savings are causal, repeatable, tokenizer-positive, and safe on held-out tenant data.
Program sequence
Small, attributable changes first. Cumulative behavior waits until each experiment earns promotion independently.
Final rollback remains unconditional. Critical-clause shielding is now the default, with an explicit benchmark-only off ablation.
Promoted defaultsStrict whitespace, safe JSON minification fallback, and repeated literal aliases.
Evaluated · parkedTokenizer-backed matrices for JSON-to-TOON and HTML-to-Markdown.
Evaluated · parkedExact approved boilerplate and aliases for classified generated wrappers only.
Needs eligible dataCombine only experiments with positive held-out savings and zero accepted hard failures.
Intentionally emptyEvidence to date
Aggregate savings across alternative arms are never presented as deployable savings.
Ten fixed cases, three repeats, and four causal conditions per profile on release 2026.07.13.205912.
Twenty-five records from one tenant corpus, representing fourteen unique texts. One repeat and six profile-plus-model arms.
default:base; tenant-specific settings were not exercised.Fifteen selected records, twelve unique baseline inputs, six profile-plus-model arms, and one repeat. Seven arm requests returned errors and are excluded from integrity rates.
Twenty-five selected records, twenty unique inputs, six profile-plus-model arms, and one repeat. All 175 configured arm records completed.
Three fixed-corpus matrices, four conditions and three repeats each: 360 records on deployment 2026.07.14.131419. A pre-fix run exposed the missing Never imply clause; the corrected v2 run is the decision source.
Decision register
Promoted permanent safety default · Run next actionable evidence candidate · Parked insufficient demand or eligible data
| Experiment | Status | What the evidence says | Next proof |
|---|---|---|---|
| Final integrity validation & rollback | Permanent default | Rejected model output never counts as savings. The two new cohorts rejected 47 unsafe candidates—35 inline-code and 12 identifier changes—while all 273 completed accepted outputs passed hard-integrity checks. | Keep unconditional on every model path and continue reporting rollback reasons separately from accepted output. |
| Critical-clause shielding | Permanent default | The corrected ablation favored shielding on every operational measure: 624 versus 486 accepted tokens saved, 6 versus 15 rollbacks, and 682 versus 741 ms p50. All categorized downstream checks passed. | Keep on by default. Retain the off profile only as a benchmark/diagnostic control; rollback remains unconditional. |
| Shielding on/off guardrail ablation | Completed | Three repeats of 10 fixed cases produced 120 records with zero errors and zero accepted integrity or downstream failures. The run also exposed and verified the repaired Never imply clause classification. | No further ablation required before release; extend the clause corpus when new policy verbs appear. |
| Expanded JSON-to-TOON | Parked | The fixed rerun did not reproduce incremental savings: baseline and experiment deterministic arms both saved 381 tokens with identical applications. The earlier 39-token tenant observation remains directional only. | Reopen only with a separate natural held-out corpus containing eligible records; do not populate safe_stack_v1. |
Safe JSON minificationjson_minify_safe | Parked | The repair is verified: zero skipped-record model-input hash mismatches. However, the experiment again applied zero minifications and added zero deterministic tokens over baseline. | Reopen only when natural traffic contains tokenizer-positive eligible JSON; the hidden model-input experiment is closed. |
| Strict prose whitespace | Parked | Eight candidate rewrites produced zero tokenizer savings and no incremental deterministic output change. | Reopen only when telemetry supplies naturally tokenizer-costly prose spacing. |
| Repeated literal aliases | Parked | No eligible repeated long URL or identifier appeared in the fixed corpus or any of the three tenant slices. | Reopen only with naturally occurring held-out examples and exact expansion checks. |
| Expanded HTML-to-Markdown | Parked | No cohort applied the transform; candidates were absent or below threshold. | Reopen only when real article/main HTML is common enough to justify a dedicated held-out corpus. |
| Tenant-approved exact boilerplate | Deferred | Discovery remains diagnostics-only. No tenant supplied an approved, versioned phrase set. | Collect at least 50 discovery records, approve a versioned exact phrase set, then evaluate separate held-out tenant data. |
| Classified duplicate-wrapper aliases | Parked | No classified generated-support wrapper appeared in either new cohort. Generic duplicate removal stayed diagnostics-only as intended. | Reopen only if production telemetry shows this explicit wrapper class is common. |
Promotion contract
A clean integrity report is necessary, but it is not the same thing as semantic task success.
Every applied record must clear absolute and relative tokenizer gates, and stage accounting must reconcile exactly. Rejected output contributes zero savings.
Protected literals, constraints, required terms, JSON, code, and structural guardrails must have zero failures on accepted output.
Negation, obligations, permissions, scope, thresholds, required formats, and entity-value relationships need explicit task evaluation—not proxy metrics alone.
Deterministic and final output hashes must be stable across at least three identical repeats, with stable skip and rollback reasons.
Discovery and evaluation records remain separate. Tenant results are reported independently, including zero-application cohorts.
Profiles are allowlisted and request-scoped. Removing a profile selection restores baseline behavior without rewriting tenant content.
Next evidence