BCJ over CM2
exe margin lever (composition unlocked)
Почему это может сжать сильнее
The arch-matched BCJ branch filter (rel->abs) was a measured win only against the old bwt-rans backend. Over CM2 it was never measurable because encode_bcj's nested encode runs with the recursion guard that disables encode_cm2. Composing them directly tests whether CM2's match/order models already subsume the rel->abs regularity, or whether turning each call target into one repeated absolute constant still buys real bits over a strong backend.
Результат проверки
замер · CUBR-bcj-cm2-exe · 2026-07-23
Validated in the shipped rail at b960b29a293ce937c14a1e91f803a153af89f477 (CLI sha256=f77878a1aea62e1771accef113df331d24d711d2011d517bcf3afd45bfbab3b5). Real CLI, byte-exact RT cmp=0, MODE_BCJ nesting MODE_CM2 wins competitive-min on ooffice: 1991681->1763460 (-11.46%, ooffice stays rank 1). CM2 does NOT subsume BCJ — the 'already learned' counter-hypothesis is refuted by ~23x the 0.5% GO threshold, and the margin is consistent with the old H-45 dense-.text result (+10.53%). Control: mozilla is a tar of PE members, bcj_detect_arch returns None, so the candidate never fires there; forcing whole-file x86 BCJ on it costs +3.78%, confirming the win is code-density bound. Zero regression: arch-detect survey over the corpus shows ooffice is the only file that can fire (sum is SPARC-detected but 38240 B < the 256 KB CM2 gate); full-corpus byte-identity verified; 293 lib tests pass.
Уроки
A pre-transform retired against a weak backend is NOT closed against a strong one — and a recursion guard can silently make a composition unreachable, so 'never measured' must be distinguished from 'measured and lost' before a branch is declared dead. Gotcha #11 says a strong entropy backend subsumes simple re-encodings; BCJ is the counter-example that marks the boundary: it does not re-encode what CM2 already models, it changes the INFORMATION by collapsing many byte forms of one call target into a single repeated constant. Nesting the new backend inside the EXISTING container kept the decoder and wire format untouched.
FH3-09
image MED16 context bias-cancellation (CALIC-class APM)
Почему это может сжать сильнее
CALIC leaves MED16 without its context error-feedback bias-cancellation component; adding a decoder-reproducible per-context running-mean correction of the true prediction error to the MED predictor, offered under competitive-min, wins exactly where it helps (x-ray) and is byte-identical where it does not (mr).
Результат проверки
замер · CUBR-fh309-image-apm-x-ray · 2026-07-23
Validated in the shipped rail at bae50473afe3f55a8682db6cae41443a6f135197 (CLI sha256=362f7cc465d3a2ce90cd5241a48341f41cfe17de650e499ce0de8138a5219c1b). Real CLI, faithful bwt-rans nested coder, byte-exact round-trip: silesia/x-ray 3771605->3699535 bytes (-1.9109%, MODE_MED16 apm=1), enc/dec rc=0, external cmp=0, archive sha256=ae494b4027db194259fbf9e3ec3359f2d4404a3a46ffa2077f29d6c6ab0a81c7. mr byte-identical (2098063, apm=0 - bias hurts mr, min() keeps plain). Image type aggregate 0.3119487161633968 -> 0.3081471588277679 (-1.22% self), widening the thinnest margin; overall 0.19128519668748242 -> 0.19105622084751853. Zero per-file regression by construction. Probe evidence: causal APM bias-cancel order-0 x-ray -2.16% / mr +11.6% (NO-GO mr); FH3-02 bit-plane NO-GO (x-ray LSB pure noise); FH3-01 MRP lower-bound x-ray +1.85% (overlaps FH3-09).
Уроки
x-ray's MED16 residual is near-order-0-incompressible (BwtRans ~7.12 b/sample vs H0 6.96), so a context bias-correction on the predictor passes through end-to-end; mr is already deeply exploited by MED16+BwtRans (its residual structure is variance, not mean) so bias-cancellation hurts it and competitive-min correctly keeps plain MED16. Bias-cancellation is the cheap CALIC sub-component; a full 2D above-row context model (GeoCM handoff) is the next, larger image lever.
CM2-EXE
strong context mixing as executable backend
Почему это может сжать сильнее
A strong CM backend can replace the missing stronger-LZ lever on executable corpora; per-file competitive selection admits it only where it wins.
Результат проверки
замер · CUBR-0064-full24-cm2-retained-min · 2026-07-22
Validated full24 strict per-file retained-min at 6eaefad7e165cd74f7d660aeb6d0828bfbe12c41: 24/24 real CLI encode+decode rc=0 and external cmp=0; 15 new rows + 9 retained verified live rows. Weighted ratios: text=0.17693132733728506, exe=0.2481901139274113, binary=0.4656539095049861, code=0.14534561821088132, database=0.08872319610228702, image=0.3119487161633968. All six types rank #1. CLI sha256=e3da2f33ef7e9f0cfb4e07c4dbe4b406ba00ab17d8b9c40bf3aeb21abd1a546a.
Уроки
CM2 beats the 7z executable aggregate on mozilla and ooffice; earlier failed executable-specific mechanisms remain historical NO-GO results.
FH-18
tar-aware DEC Alpha ECOFF BCJ per member (mozilla, exe)
Почему это может сжать сильнее
mozilla (89% exe-байтов) — tar, члены не ELF/PE: крупные .so несут ECOFF DEC Alpha magic 0x0183 (bytes 83 01 15 00), BR/BSR branch density ~2.03% + conditional ~8.26% (evidence BUILD-:19). Гипотеза: tar-aware проход по членам, Alpha/ECOFF BCJ-фильтр per member, с ЧЕСТНО charged tar framing (заголовки/выравнивание — decoder branch по Gotcha #6/#7). Floor (операторский, 2026-07-16): target mozilla RT=OK/cmp=0 и ratio < min(7z,xz) на mozilla; затем full24 exe aggregate < 0.2748738268008853 (7z type leader). Tie/в пределах округления = NO-GO. Очередь: после FH-01 StateMap; run владеет C1, BUILD-:19 — только evidence/design. Переиспользуемый блок: ooffice BCJ+CM win из FH-07 (2121874 B).
Результат проверки
замер · CUBR-0046 exe per-type queue · 2026-07-16
target-only dev-ai 2026-07-16 (full24 не запускался — операторский гейт провален на target): mozilla candidate tar-Alpha-BCJ 15924339 B (0.310897886938974) ХУЖЕ собственного baseline 15788540 B (0.308246623225710) на +135799 B; гейт требовал < min(7z 0.260534184763595, xz 0.261150227409036). RT=OK cmp=0 — механизм корректен (525-member ustar, ECOFF 0x0183, BR/BSR→absolute, member-map ре-деривируется, charged только 10-Б контейнер), но переписывание BR/BSR не создаёт эксплуатируемой избыточности для нижележащего кодера. Структурный вывод: разрыв Cubrim 0.3082 vs 7z/xz 0.2605 на mozilla — это класс силы LZ-бекенда (LZMA match modeling), НЕ отсутствующий BCJ (у 7z/xz Alpha-BCJ нет, а они лидируют). Exe #1 через BCJ-семейство закрыт; см. Gotcha #8/#9/#10 о пределах атаки на LZ-бекенд.
NEW-24
Fast-CM: универсальный context-mixing top-rail с бюджетом времени (профиль --max)
ЗАКРЫТА
Почему это может сжать сильнее
Ветка MODE_CM: бинарная декомпозиция байта (дерево из 8 бит) → модели: match model (хэш длинных контекстов), order-1..3 (хэшированные счётчики состояний), опционально word-model для text и sparse-контексты для binary → logistic mixing (16–64 весов, выбор набора по фиче-классу от NEW-21) → APM/SSE финальная калибровка → бинарный арифметик/rANS. Включение: профиль --max и/или бюджет NEW-22.5. Цели: **enwik8 ≤0.218** (сейчас 0.2622, ppmd 0.2240, gap +17.0%), **webster ≤0.155** (0.2105, ppmd 0.15...
Результат проверки
замер · 2026-08-11
не реализовано — GO (флагман): MODE_CM (logistic mixing+SSE) = реализация NEW-01/H-61 backend-lever | 2026-08-09: decode-attribution characterisation preregistered (PR #51, main d212c1c) and running on dev-ai — groundwork before any Fast-CM lever. | 2026-08-09 MEASURED (characterisation, PR #54): decode budget on CM2 cells = ~50% probe-load latency (predict_bit), ~33% Ctr::upd write-back, ~7% mixer/APM update, 0.45-0.59% range coder; IPC 1.46-1.89, single-threaded; max->web speedup entirely reduced memory stalling. Compound lever = fewer models per bit (time-budgeted selection) — SIMD/rANS directions dead at <=1.005-1.1x. x-ray decode 98% geocm, 0% CM2.
[EXTENSION 2026-08-09 · CUBR-DECODE-ATTRIB-G2-VALID] Valid G2 decode-time characterization completed once on dev-ai with pin 0-15 and landed in PR #66. This is the canonical characterization pointer. Prior PR #54 G0 and PR #58 tier artifacts used the invalid 16-19 path and remain exploratory-only; none of their measurements, verdicts, tier ranking, or lever selection is adopted here. G2 verdicts: P1 and P4 SUPPORTED; P2, P3, and P5 INDETERMINATE. No candidate, benchmark throughput, measurement row, evaluation, or density/speed trade was established. Evidence: documentation/ephemeral/research/CUBR-DECODE-ATTRIB-RESULTS-20260809.md and documentation/ephemeral/research/CUBR-DECODE-ATTRIB-G2-RESULTS-20260809/. NEW-24 remains in_progress. | 2026-08-11 MEASURED (PR #101, main 94c9555): F12 2.24x / M8 3.37x decode on dickens (floors 1.5x/2.0x, map 1.81x/2.35x — above-map = working-set shrink on the stall term); osdb M8S 2.77x. Density whole-file: F12 dickens +3.58%/osdb +8.85% (P-B thresholds refuted, lead survives all five); M8 osdb +9.51% lead lost, M8S +6.50% recovers; samba +1.82/+4.57, xml +4.14/+8.62, enwik8-head +1.73/+6.08. Wire: 1 decoder branch, 0-byte header charge, old parse fails closed O(1). Defaults byte-identical; preset adoption = future preregistered campaign. Record: CUBR-NEW24-TIERS-20260811-results.md
Вердикт консилиума
GO (флагманский MODE_CM = H-61 lever = реализация NEW-01)
Далее: Реализовать MODE_CM: бинарная декомпозиция байта → модели {match, order-1..3 хэш-счётчики, word для text, sparse для binary} → logistic mixing (16-64 входа) → SSE. Это КОНКРЕТНАЯ реализация NEW-01/H-61 (флагманский бэкенд). Высокое усилие/риск. Follow-up в backlog (единый CM-трек). Критерий: сокращение gap к ppmd с +17..33% до минимума на text/code без регрессии binary.
Уроки
Оба вендора GO. Урок: после исчерпания трансформ (grid 6/6) lever ОБЯЗАТЕЛЬНО смещается в бэкенд контекстного моделирования (#1, H-61); logistic mixing+SSE — прямой путь к ppmd-паритету. Это и есть реализация NEW-01.