Bottom line: Rising token costs for frontier models like Opus or current GPT versions are pushing enterprises toward dynamic model routing, as demonstrated by Glean with a cost structure it says is four times cheaper than Claude Code.
Rising prices for top-tier models such as Opus or current GPT versions, along with the growing capabilities of open models like Kimi K3 and Qwen3.8-Max, are driving enterprises to dynamically manage model access instead of committing to a single provider. This trend is visible both in Stripe’s acquisition of OpenRouter for over 7 billion US dollars and in the enterprise business of providers like Glean.
Glean, co-founded and led by former Google Distinguished Engineer Arvind Jain, positions itself as an AI platform for large organizations. Following a Series F funding round of 150 million US dollars last June, the company was most recently valued at 7.2 billion US dollars. This year, Glean reportedly reached 300 million US dollars in annual recurring revenue (ARR) — a tripling within 15 months. A central component of the product is selecting which model is used for which task — or whether an LLM is needed at all. Jain cites simple arithmetic operations as an example, for which users sometimes use LLMs instead of a calculator.
According to Jain, Glean offers three levels of model selection: employees can explicitly choose a model, administrators can restrict models or set usage limits, and an automatic mode dynamically selects the appropriate model for each task. Jain states that customers predominantly opt for the automatic mode — for economic reasons. Co-founder and engineering lead Tony Gentilcore put Glean’s cost efficiency relative to Claude Code at a factor of 4: an average of 0.45 US dollars per task compared to 1.84 US dollars for Claude Cowork, attributed to Glean’s harness and routing capabilities.
For CTOs, the cost trajectory of frontier models is the actual driver of this development. Jain describes how current top-tier models like Opus or the latest GPT versions are sometimes two to four times more expensive per token than their predecessors, while simultaneously being used for significantly longer and more complex tasks. In sum, costs per user could increase tenfold to twentyfold compared to the previous year. While individual users are well served with subscriptions ranging from 20 to 200 US dollars per month, this model does not scale in enterprises with thousands of employees without control mechanisms.
Glean explicitly positions itself as a meta-harness spanning multiple model providers, describing its product, according to Jain, as a “superset of ChatGPT, Claude, Gemini and Grok” within a unified user interface. Since the third generation of the Glean Assistant in September, agents have also played a growing role in the system. Glean cites reference customers including Zillow, with an 80 percent adoption rate among 7,000 employees, as well as Booking.com.
For CTOs planning or evaluating their own AI infrastructure, the report offers a benchmark for assessing routing layers: cost savings arise less from a single cheaper model and more from an architecture that distributes tasks according to complexity across different models or even classic systems without an LLM. The combination of explicit user choice, administrative guardrails, and automatic routing forms a pattern that can also be observed beyond Glean as a reference architecture for enterprise AI deployments.
Source: www.latent.space · Published August 18, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.