The problem

Receipts and invoices are deceptively hard. The text is often legible to OCR while the structure is not: which number is the total, which line is tax, which column belongs to which item. Pure OCR reads characters but has no model of a document; pure LLM reasoning is expensive and hallucinates digits.

What I built

A hybrid system that gives each component the job it is actually good at:

  • Native OCR for character recognition, where it is fast, cheap, and more accurate than a language model.
  • LLM-based parsing for layout reasoning and field assignment: deciding what the extracted text means.

It reached 92% accuracy on internal benchmarks, and was optimized for deployment on edge devices rather than assuming a server round trip.

Also built here

An LLM-powered chatbot that personalizes responses using hybrid RAG retrieval, combining semantic and keyword search, so answers stay grounded in the specific user’s information rather than general knowledge.