About
A firewall-disciplined toolkit and methodology for the computational study of the Voynich Manuscript (Beinecke MS 408) and other ground-truth-free corpora.
Computational Voynich work has a long record of confident, mutually incompatible "solutions," because with no ground truth "it looks like X" is nearly unfalsifiable. This project inverts the order: validate first, claim second. Four coupled disciplines make that concrete:
- Harness — every method must separate real language from matched structured-meaningless controls and ciphers before it touches the manuscript.
- Firewall — every number comes from deterministic, versioned code; nothing is estimated or recalled.
- Evidence grading — every claim carries an A–D grade and never gets upgraded to look stronger.
- Adversarial refutation — every strong claim is attacked by an independent clean-context reviewer before it stands.
The statistical evaluator is cold and reproducible; the accompanying refutation protocol uses a fallible language model and is a discipline, not an oracle. Read the methodology and the limits.
Everything is open source under Apache-2.0. Source, issues, and contribution guidelines are on GitHub.
Why this project
Why MS 408?
This program sits where a few long-running threads collide: a computer scientist with a natural-language-processing and human–computer-interaction background; a career as a CTO, systems designer, and builder; a parallel, self-directed curiosity about AI; and, alongside it, an amateur's fascination with anthropology and human potential — how people, and now machines, make meaning.
The Voynich Manuscript is the natural meeting point for all of them. It is the ultimate ground-truth-free corpus: six centuries old, richly structured, and stubbornly unread. The specific spark was a question about method and models — what happens when a frontier model (Claude Fable 5) is pointed at the manuscript, but held inside a strict firewall discipline built to stop it (and its operator) from fooling itself? Would the two collide, or reconcile? This toolkit is the result: not a solution, but a cold, honest way to see how far rigorous statistics can actually take you — and exactly where they stop.
About the author
Tim Walsh
Tim is a computer scientist, systems designer, and technical founder based in Scottsdale, Arizona — co-founder/CTO across several startups, with a background spanning full-stack engineering, enterprise architecture, human–computer interaction, and natural-language processing (B.Eng. in Computer Science, University of Delaware, with supplementary study in psychology and cognitive science). MS408 is an independent research project run under DireLabs, and reflects a longstanding interest in applying rigorous method to genuinely open questions.
Personal site · LinkedIn · GitHub · an early ACM Student Research Competition paper (2010)