MODULAWAI
The retrieval layer under ModulawAI's legal research. We built the engine that turns statutes and case law into something a model can search and cite.
- Role
- Ingestion, parsing, embedding, retrieval
- Corpus
- Statutes and case law
- Status
- In service

Scope
BEESTUDIOX built the ingestion and embedding engine. The wider ModulawAI platform, including its matter management, billing and client tooling, is not our work and is not described here.
Legal sources are not a tidy corpus.
A legal research product is only as good as what sits underneath it, and the sources it has to read are the problem. Every publisher formats differently. Structure is implied by typography rather than marked up. Citation conventions are inconsistent, and a single judgment can run to hundreds of pages.
Drop that into a generic retrieval pipeline and it will happily return a paragraph of text with no idea which section of which Act it came from. That answer is worse than no answer, because it looks usable.
The engine had to preserve the structure the law is written in, end to end, from ingestion through to the passage that comes back from a query.
What we built
FOUR STAGES, ONE PIPELINE.
Ingestion
Sources publish in whatever shape suits them, change without notice, and republish the same judgment under a different reference.
- One pipeline per source, isolated so a broken source cannot stall the rest
- Normalisation into a single internal document shape
- Change detection, so a re-publish updates rather than duplicates
- Failures surfaced as failures, never as a silently empty result
Parsing and structure
This is where the work is. A statute and a judgment are not prose, they are hierarchies, and most of the value sits in that hierarchy.
- Statutory hierarchy recovered: part, section, subsection, paragraph
- Judgment structure recovered, including paragraph numbering
- Citation references extracted and resolved to what they point at
- Scanned and malformed documents routed down a separate path
Chunking and embedding
Generic fixed-length chunking destroys the one property legal retrieval needs: a result has to be something a lawyer can cite.
- Chunk boundaries follow the document's own structure, not a token count
- Every chunk carries the reference needed to cite it
- Embedding and index strategy chosen against cost at corpus scale
- Re-indexing treated as routine, because corpus and model both move
Retrieval
The platform asks the question. This layer returns passages that are relevant and quotable, with the authority still attached.
- Hybrid retrieval rather than vector similarity alone
- Filtering by jurisdiction, court and date
- Results carry their citation, so the layer above can show its working
- Cost per query bounded, because research is never a one-shot call
DECISIONS THAT HELD UP.
A chunk has to be citable
The constraint that drove everything else. If a retrieved passage cannot be pointed back to a section or a numbered paragraph, a lawyer cannot use it, however good the similarity score was.
Parsing failures are loud
A pipeline that silently drops a malformed judgment looks healthy right up until someone searches for that judgment. Failures are recorded against the document and visible, not swallowed.
Re-indexing is a routine operation
Embedding models change and the corpus grows. Treating a full re-index as an emergency guarantees the index falls behind, so it was built to be run on purpose rather than in a panic.
Cost per document sets the ceiling
Whether a corpus can keep growing is an economics question before it is an engineering one. Ingestion and embedding cost was a design input rather than something measured afterwards.
Where it stands
The engine is in service and still being extended as the corpus grows. ModulawAI is the product it feeds.
Visit ModulawAI →GOT A CORPUS NOTHING CAN READ?
Ingestion and retrieval over messy, high-stakes documents is the work we are best at. Tell us what your sources look like.
Start a project