
Applied AI · 2025–present
TracEx
Every number pulled out of a paper carries the exact sentence it came from, and a grade for how well that check holds.
The short version
Measured values in materials-science papers are buried in prose: finding one parameter meant reading the PDF and copying the number by hand. A plain LLM call or a generic AI tool returns results full of hallucinations. What was missing was extraction where every number ties back to the exact sentence it came from.
Pythondoclingpymatgenquantulum3litellmpytestLLM tool useJSON Schema
How
- Stamped provenance before extraction. docling parses the PDF, and every sentence and table cell gets a hierarchical ID before any LLM runs, so every claim ties to a fixed text anchor.
- Ran deterministic extraction first: regex, pymatgen for chemical-formula validation and quantulum3, extended with a condensed-matter unit registry, for numbers and units. Only then comes one LLM call per paper, with forced tool use against a JSON schema and a hard call budget.
- Kept the model provider swappable through litellm, across Anthropic, OpenAI and Gemini.
- Graded every value deterministically afterwards: the LLM's cited sentence IDs are re-checked against the original text, with substring alias matching for sample names and word-boundary regex for numbers, giving grade A, B or C.
- Wrote one Obsidian markdown note per sample, with a measurement table carrying source links and grades, plus JSON artefacts for each stage. Tested with pytest.
The same stamp-then-verify approach underlies the lab's literature retrieval.