Evidence ledger · Updated July 15, 2026

Compression experiments, with receipts.

A living record of what we tried, what actually saved tokens, what stayed intact, and what still needs proof. Experiments graduate only when savings are causal, repeatable, tokenizer-positive, and safe on held-out tenant data.

10tracked experiments and safety decisions
7completed evidence cohorts used in the register
0savings transforms promoted to safe_stack_v1
2permanent safety defaults: shielding and rollback

Program sequence

Current phases

Small, attributable changes first. Cumulative behavior waits until each experiment earns promotion independently.

0

Measurement & integrity

Final rollback remains unconditional. Critical-clause shielding is now the default, with an explicit benchmark-only off ablation.

Promoted defaults
1

Existing safe features

Strict whitespace, safe JSON minification fallback, and repeated literal aliases.

Evaluated · parked
2

Threshold expansion

Tokenizer-backed matrices for JSON-to-TOON and HTML-to-Markdown.

Evaluated · parked
3

Tenant-specific structure

Exact approved boilerplate and aliases for classified generated wrappers only.

Needs eligible data
4

Safe stack

Combine only experiments with positive held-out savings and zero accepted hard failures.

Intentionally empty

Evidence to date

Benchmark cohorts

Aggregate savings across alternative arms are never presented as deployable savings.

Release matrix

Fixed safety corpus

No promotions

Ten fixed cases, three repeats, and four causal conditions per profile on release 2026.07.13.205912.

840full-matrix records
0accepted hard failures
0incremental deterministic tokens
  • Deterministic baseline and experiment arms were identical.
  • Experiment-plus-model accepted 10.5% total savings after rollback.
  • No savings experiment met the positive held-out criterion.
Tenant subset

Delivery Tower prompt slice

Directional only

Twenty-five records from one tenant corpus, representing fourteen unique texts. One repeat and six profile-plus-model arms.

15.9%one-arm model savings
0deterministic applications
100%hard protected-span retention
  • Savings came entirely from LLMLingua, not the named deterministic profiles.
  • The service used default:base; tenant-specific settings were not exercised.
  • Constraint and required-term coverage were zero, so semantic acceptance remains open.
Tenant 1

Large structured-prompt slice

Positive TOON signal

Fifteen selected records, twelve unique baseline inputs, six profile-plus-model arms, and one repeat. Seven arm requests returned errors and are excluded from integrity rates.

39incremental deterministic tokens
1matched TOON application
0accepted hard failures
  • Expanded TOON changed one of fourteen completed matched records: a positive 0.01% signal, not promotion evidence.
  • The other five profiles produced zero incremental deterministic savings.
  • Twenty-nine unsafe inline-code model candidates were rejected; constraint and required-term coverage remained zero.
Tenant 2

Structured workflow slice

No incremental profile savings

Twenty-five selected records, twenty unique inputs, six profile-plus-model arms, and one repeat. All 175 configured arm records completed.

0incremental deterministic tokens
18unsafe model candidates rejected
100%accepted hard-integrity pass rate
  • Baseline TOON already saved 1,428 tokens on two records; every experiment profile matched that deterministic output exactly.
  • Each experiment-plus-model arm saved 40 model tokens, but there was no model-only arm for causal attribution.
  • The run used one repeat and crossed an application deployment boundary, so it cannot support promotion.
Focused release matrix · July 15

Safety default and final transform decisions

Safety confirmed

Three fixed-corpus matrices, four conditions and three repeats each: 360 records on deployment 2026.07.14.131419. A pre-fix run exposed the missing Never imply clause; the corrected v2 run is the decision source.

0errors or accepted hard failures
9fewer rollbacks with shielding on
138more accepted tokens saved with shielding on
  • Shielding on: 624 accepted tokens saved, 6 rollbacks, 682 ms p50. Shielding off: 486 tokens, 15 rollbacks, 741 ms p50.
  • All categorized relationship, negation, permission, and required-format checks passed after the detector repair.
  • TOON and JSON experiment deterministic arms matched baseline at 381 tokens saved; neither added an application or token.

Decision register

Experiment results

Promoted permanent safety default · Run next actionable evidence candidate · Parked insufficient demand or eligible data

ExperimentStatusWhat the evidence saysNext proof
Final integrity validation & rollbackPermanent defaultRejected model output never counts as savings. The two new cohorts rejected 47 unsafe candidates—35 inline-code and 12 identifier changes—while all 273 completed accepted outputs passed hard-integrity checks.Keep unconditional on every model path and continue reporting rollback reasons separately from accepted output.
Critical-clause shieldingPermanent defaultThe corrected ablation favored shielding on every operational measure: 624 versus 486 accepted tokens saved, 6 versus 15 rollbacks, and 682 versus 741 ms p50. All categorized downstream checks passed.Keep on by default. Retain the off profile only as a benchmark/diagnostic control; rollback remains unconditional.
Shielding on/off guardrail ablationCompletedThree repeats of 10 fixed cases produced 120 records with zero errors and zero accepted integrity or downstream failures. The run also exposed and verified the repaired Never imply clause classification.No further ablation required before release; extend the clause corpus when new policy verbs appear.
Expanded JSON-to-TOONParkedThe fixed rerun did not reproduce incremental savings: baseline and experiment deterministic arms both saved 381 tokens with identical applications. The earlier 39-token tenant observation remains directional only.Reopen only with a separate natural held-out corpus containing eligible records; do not populate safe_stack_v1.
Safe JSON minification
json_minify_safe
ParkedThe repair is verified: zero skipped-record model-input hash mismatches. However, the experiment again applied zero minifications and added zero deterministic tokens over baseline.Reopen only when natural traffic contains tokenizer-positive eligible JSON; the hidden model-input experiment is closed.
Strict prose whitespaceParkedEight candidate rewrites produced zero tokenizer savings and no incremental deterministic output change.Reopen only when telemetry supplies naturally tokenizer-costly prose spacing.
Repeated literal aliasesParkedNo eligible repeated long URL or identifier appeared in the fixed corpus or any of the three tenant slices.Reopen only with naturally occurring held-out examples and exact expansion checks.
Expanded HTML-to-MarkdownParkedNo cohort applied the transform; candidates were absent or below threshold.Reopen only when real article/main HTML is common enough to justify a dedicated held-out corpus.
Tenant-approved exact boilerplateDeferredDiscovery remains diagnostics-only. No tenant supplied an approved, versioned phrase set.Collect at least 50 discovery records, approve a versioned exact phrase set, then evaluate separate held-out tenant data.
Classified duplicate-wrapper aliasesParkedNo classified generated-support wrapper appeared in either new cohort. Generic duplicate removal stayed diagnostics-only as intended.Reopen only if production telemetry shows this explicit wrapper class is common.

Promotion contract

What “safe” means here

A clean integrity report is necessary, but it is not the same thing as semantic task success.

Causal savings

Every applied record must clear absolute and relative tokenizer gates, and stage accounting must reconcile exactly. Rejected output contributes zero savings.

Hard integrity

Protected literals, constraints, required terms, JSON, code, and structural guardrails must have zero failures on accepted output.

Downstream fidelity

Negation, obligations, permissions, scope, thresholds, required formats, and entity-value relationships need explicit task evaluation—not proxy metrics alone.

Repeatability

Deterministic and final output hashes must be stable across at least three identical repeats, with stable skip and rollback reasons.

Held-out evidence

Discovery and evaluation records remain separate. Tenant results are reported independently, including zero-application cohorts.

Reversible rollout

Profiles are allowlisted and request-scoped. Removing a profile selection restores baseline behavior without rewriting tenant content.

Next evidence

Required shape of the next run

Run all four causal arms

1 baseline deterministic 2 experiment deterministic 3 baseline model-only · deterministic off 4 experiment + model-force

Use identical prompt order, tenant profile, tokenizer, model revision, aggressiveness, and at least three repeats.

Close the current evidence gaps

  • Pass the actual versioned tenant profile.
  • Verified: categorized relationship, negation, permission, and required-format checks passed in the corrected focused run.
  • Verified: skipped-JSON model-input neutrality had zero matched hash differences.
  • Verified: the focused runner completed 360 records with zero harness/API errors.
  • Do not run another savings-transform matrix until a natural held-out sample contains eligible, tokenizer-positive records.
  • Include deliberately eligible records plus a held-out natural sample.