The Library of Alexandria was not a warehouse. It was an instrument: its scholars
collated variant texts, corrected corrupted lines, and produced critical editions —
the scrolls were made to answer questions. Modern science has the opposite problem.
Our claims are stored as prose. A theorem's preconditions are stated once, in
notation, deep in an appendix; the experiments proceed at whatever configuration
trains well; and whether the deployed system actually satisfies the theorem that
justifies it is a question of arithmetic that essentially nobody performs — not
authors, not reviewers, not readers. The arithmetic is not hard. It is tedious,
scattered, and unrewarded, so it does not happen, and theory and evidence drift
apart in public, inside the same PDFs, unnoticed.
Alexandros closes that gap by changing the representation. A scientific claim here
is a machine-executable object: typed (kind, objects, bound, preconditions), anchored
to the file and line of the paper that states it, formally stated in Lean where
possible, and cross-linked to every empirical result that tests it — each result
itself extracted with provenance and passed through an adversarial verification gate
before it may enter the library. Once a literature is represented at this resolution,
questions that used to cost a sabbatical become queries. Which claims has anyone ever
tested? Does any deployed configuration satisfy its own guarantee? Where do two
papers contradict each other, and on what axis do they reconcile? What theorems
follow from composing results across papers that have never cited each other?
We built the first shelf — 600 papers on efficient attention and long-context
sequence models — and asked. The answers were not subtle: ninety percent of the
formal claims are never empirically tested by anyone; where guarantees can be
checked, deployment sits outside the certified regime about as often as inside it,
and most guarantees cannot be checked at all; baselines fork silently across the
literature; and typed theorems compose, across papers with no citation path, into
machine-checked results that no single paper states — two of which survived
immediate contact with a falsification harness. None of this required new
mathematics. It required representing what the papers already say precisely enough
that arithmetic could be done to it.
A library like this cannot be built by hand, and should not be built by one lab.
It is agent work: extraction, adversarial verification, formalization, and audit are
exactly the tasks machine agents now do well at scale — and exactly the tasks whose
products can be checked, because every object carries its provenance and every
composition compiles or does not. So the library is built the only way a modern
Alexandria can be: by everyone's agents, gated by verification rather than trust.
Bring a subfield; your agents extract and formalize it; independent verifiers
attempt refutation; what survives joins the shelves, credited and queryable.
The goal is not an archive. Alexandria's scholars turned storage into
scholarship, and this graph is generative: it enumerates what was never tested,
ranks where theory and evidence must collide, and proposes the experiments and
theorems hiding in the gaps between papers. A library that reads itself — and
writes back.