Work done on bodies of knowledge — finding what a corpus implies but never states, compressing it to what matters, and building the corpus in the first place. Methods validated on real material, not demonstrations.
Coming soon
A command and control centre that relays real-time data and intelligence to a distributed team working to eliminate the threat of global terrorism.
The same machinery as the work below — a corpus, retrieval, and a model that reads in context — pointed at signals arriving now rather than at a library that sits still.
Feeds, reports and field data land continuously and are embedded as they come, not in a nightly batch.
Vector search puts each signal beside everything already known about it, and a model reads it in that context rather than alone.
Webhooks fire per use case, so a signal reaches the team that can act on it — and does not reach anyone else.
Stack — web application · LLM in the loop · vector database · webhook routing per use case
📄 read the V1 design doc — scope, components, key decisions and open questions, one page
Nothing in this section is built. It is here because it is the direction, not because it ships.
What we do
Each one takes a body of knowledge — a library, a documentation set, a research programme, a shelf of recorded lectures — and returns something you could not get by reading it faster.
Derive what a coherent system implies but never states. Not what the material says — what it requires, and nobody wrote down. Detailed below.
A 750-page textbook returned as two and a half pages worth reading. Whole sources are read, never abstracts or summaries. Papers, books, lectures and recorded talks.
The inverse of gap analysis. Pull one coherent system out of many sources that disagree, and find where two fields discovered the same result without citing each other.
Build the searchable body of knowledge to begin with — ingestion, chunking, embedding, retrieval — so the work above has something to run against.
The flagship · KEGA
Every coherent body of knowledge implies content it never states. KEGA reads the structure surrounding a gap and derives what the gap must contain — without new data, and without waiting for someone to discover it.
Traditional research treats a gap as a problem of discovery — you find the missing piece or you don't. KEGA treats it as a problem of derivation. A system's surviving structure bounds what its missing content can logically be, and the tighter the coherence, the narrower those bounds.
Derive gap content from the internal logic of a formal system.
Output — a verifiable claimGiven a theory, derive the experiments that would verify or break it.
Output — constrained, falsifiableDerive what a narrative or creative world implies but never states.
Output — a plausible claimThe output is never a claim to be the original. It is a claim to be implied by the system's own logic — the doppelgänger, not the original. Same structural soul, different vessel.
Validation
One modern and mathematical, one classical and philosophical, 2,400 years apart. Both produced output structurally indistinguishable from the verified content around it.
Gap analysis on the textbook's derivation chain pointed at a missing link in the generative-model sequence. Working only from what the book contained, the method derived the shape of the theory that belonged there.
It matches diffusion-model theory, published separately as DDPM in 2020.
Twelve chapters survive; two were lost. From the surviving framework the method reconstructed both, then derived four further chapters the system implies but never had.
轉丸 · 胠亂
recovered · and implied
知止 无形 传道 归虚
Convergence across two unrelated domains is the finding, not either result alone.
Where gap analysis pays
Any organisation that maintains a corpus — documentation, a curriculum, a research programme, a patent portfolio — is sitting on structure it has never interrogated. The question KEGA answers is not what does this say. It is what does this imply that nobody has written down.
Stage
Worth being exact about where this is, because the gap between the two is the whole investment case.
The framework, the paper, and two independent validations across unrelated domains. A working corpus pipeline that reads full books rather than abstracts.
A forecasting run: take a frontier field, derive what its next breakthrough must look like, publish it timestamped, and wait for the real one to land.
Runway to turn a method into a service other people can point at their own corpus without a researcher sitting beside them.
Get in touch
The four above are what has been built and proven. If you have a body of knowledge and a question that doesn't fit any of them, that is worth a conversation too. Leave an address and what you'd want read. No list, no newsletter — a direct reply.
Or write directly — davidkwokhochan@gmail.com
Nothing is shared or published. Submissions reach the author only.