Architecture Vision
Which AI investments make sense today
LLMs are changing faster than enterprise roadmaps. The durable investment is the capability to delegate agentic workloads safely, evaluate AI output reliably, and route each workload to the right model.
From assistance to delegated workloads
Until 2025, LLMs mainly assisted with individual steps. Since autumn 2025, strong models have been able to complete clearly described software tasks over several hours, including tool calls, tests, and corrections. An independent benchmark shows the scale of the change: GPT-4 handled tasks measured in minutes, while Claude Opus 4.5 handled tasks lasting more than five hours by the end of 2025.
The same way of working is now reaching research and other office workflows. For enterprises, the specific AI solution matters less than the operating boundary: which agentic workloads may be delegated, which systems and data may agents access, and which criteria determine whether their work is correct for the business?[1]
Frontier capability no longer belongs to frontier labs alone
OpenAI released GPT-5.4 on 5 March 2026. Just over five months later, DeepSeek V4 Pro 0813 reached the same Intelligence Index score of 53. The estimated cost of the compared agentic workload was 1.10 US dollars for GPT-5.4 and 0.25 dollars for DeepSeek. A capability level reached only by a proprietary frontier LLM in March was available as open weights within one planning cycle.
A tie on the index does not, however, make the models interchangeable. This is because the index aggregates nine evaluation areas; a model can be optimised for known benchmarks and still behave differently on proprietary code, German contracts, or domain-sensitive decisions. The comparison is therefore a market signal, not a selection decision. Selection requires enterprise evaluations and staff who understand the business process well enough to interpret real tests.[2]
Intelligence per dollar is now measurable
Across six benchmarks, the same measured level of model performance became 9 to 900 times less expensive within a year depending on the task. That range shows that there is no universal rate at which AI capability becomes cheaper. A cost comparison applies only to a specific workload, a defined quality level, and a stated date.
Enterprises should therefore compare the total cost to business acceptance, not token prices: in software, until a change has passed tests and review; in knowledge work, until a subject-matter expert has reviewed and adopted the result. DeepSWE provides measured success, time, and cost for software tasks under one agent setup. Formula v1 divides attempt time and cost by first-run success, making the cost of unreliable output visible. Its weighting is an editorial starting point, not a business-value model. Enterprises must repeat the calculation with their own tasks, acceptance rules, review effort, and failure costs.[3][4][5]
Page snapshot9 October 2026113 tasks
Bottom left: faster and cheaper. Diamonds sit on the Pareto frontier.
2 off-scale outliers remain available in the full table.
Models28/28
Leading LLMGPT-6 Astra