Syntheia Introduces Innovative Method to Reduce Token Costs in Legal AI Applications

Syntheia, led by founder Horace Wu, has developed a novel approach to contract analysis that significantly reduces token costs associated with AI applications in legal contexts. This method focuses on how legal documents are structured and processed before AI analysis, rather than merely switching to less expensive language models.
The rising costs of tokens in legal technology have prompted a broader discussion within the industry. Syntheia's findings suggest that the expense incurred when an AI model reads a document often outweighs the cost of reasoning over the answers it generates. The company has highlighted that, in their testing, the length of the final answers produced remained consistent across different methodologies, while the volume of text processed by the AI varied significantly.
Through their research, Syntheia's team tested two structured retrieval methodologies designed for transactional legal texts. These methodologies were evaluated against full document injection using a benchmark comprising real credit facility agreements, limited partnership agreements, and share purchase agreements. The results indicate that by honing in on the most relevant information, substantial reductions in token usage can be achieved.
The research, initiated two months prior, employed Claude 4.6 as the primary engine for question and answer tasks, specifically targeting transactional documents. The approach aims to assist law firms and generative AI companies in efficiently analysing contracts to address specific queries.
To maintain accuracy while reducing token consumption, Syntheia's method involves minimising unnecessary tokens within the context window and eliminating the need for multiple passes over documents to extract defined terms. This is accomplished through a structured document framework that establishes connections between clauses and definitions, ensuring that accuracy remains intact even as token usage decreases.
Syntheia's approach to retrieval, referred to as RAG (retrieval augmented generation), diverges from conventional methods that rely on semantic similarity. Instead, the company employs a reasoning based retrieval system, allowing the AI to determine which clauses are necessary to answer a given question based on a concise index of the document.
The token savings achieved through this method have been benchmarked against both full document injection and vector RAG approaches, revealing that while vector RAG provides a satisfactory level of accuracy, Syntheia's method offers significant savings without compromising quality. However, the research is still in its early stages, with plans for further exploration and development on the horizon.
When considering how law firms should manage token costs, the advice varies. Firms that have partnered with platforms like Harvey and Legora may benefit from all you can eat pricing models, while others that rely on Claude or similar models must focus on minimising their token expenditure while ensuring high quality outputs. The question of whether to pass these costs onto clients remains a topic for commercial discussion.
As the legal technology landscape evolves, demonstrating token savings is likely to become a key selling point for AI solutions. The ability to ensure that AI generated outputs are both accurate and cost effective will be crucial for law firms looking to leverage technology for enhanced legal outcomes.