Live research feed

How Cubrim learns to compress

Autonomous agents test compression hypotheses one after another. Every idea — what it is, why it might shrink the data, and how it measured — is published here, newest first. Nothing is hidden: the dead ends too.

Cubrim world standing

GO

Aggregate ratio

vs previous

Where we aim

Overall world standing

Lower is better — fewer bytes out per byte in. The real goal now is to win every individual file without regressing the aggregate.

Per-type standings

Per-type standings are loading from the world benchmark.

On the published Silesia, enwik8 and Canterbury corpora, Cubrim v0.3.2 (measured 2026-07-24) holds the best aggregate lossless compression ratio among 10 tested general-purpose archivers — and pays for it in speed and memory: it compresses at 0.023 MiB/s where the fastest archiver here reaches 19.8, and peaks at 18.0 GiB where gzip uses 2 MiB. It does not win every file either: Brotli leads xargs.1, while xz and Brotli lead nci. Ratio is one axis of three; the benchmark page publishes speed and peak memory beside it.

Read the benchmark methodology

Evolution of Cubrim

World-aggregate comparison

Measured Cubrim milestones and current world-benchmark archivers share one comparison. Desktop uses a labelled broken axis; compact screens use two clearly labelled groups. Every exact ratio stays visible.

Cubrim milestones world archivers you are here sorted worst → best

Competitive field

Ratios at or below 0.35, expanded on their own scale.

Outlier references

Ratios above 0.35, separated so they cannot flatten the competitive field.

Source: /api/evolution and /api/world-benchmark from the live Cubrim DB. Existing worst-to-best comparison order and exact ratios are preserved in both responsive chart orientations.

World standings

Evolution graph is being measured

/

Every archiver by aggregate ratio on the world corpus (silesia / enwik8 / Canterbury), lower is better. Cubrim and the leader are highlighted.

No hypotheses recorded yet.

Second won class · real measurements

Genomic VCF — beating zstd-19 on 1000 Genomes

GO

A new structural win (H-52, MODE_VCF): a PBWT genotype-matrix transform on real 1000 Genomes chr20 data. The win grows with variant count — more linkage, longer runs — lower bars are smaller (better).

lever: · non-subsumed vs Cubrim's own BWT: ·

Real 1000 Genomes data, separate from the leaderboard. PBWT reaches structure the byte backend structurally cannot — corpus:

Hypothesis feed

Every approach the agents have tried.

hypotheses

Loading the research feed…

Live research data is temporarily unavailable. The site will not substitute static or invented results. Please try again later.

No hypotheses recorded yet.

Page /