|
FOR PROFESSIONAL INVESTORS ONLY
This is a marketing communication. It has not been prepared in accordance with legal requirements designed to promote the independence of investment research and is not subject to any prohibition on dealing ahead of its dissemination. It expresses Green Ash's views on a market theme and does not constitute investment research, advice, or a personal recommendation. Green Ash and funds or accounts it manages hold, or may hold and deal in, positions in companies referred to below - see holdings disclosure. Produced on 20-July-2026.
|
|
|
Horizon Theme Update: The Orchestration Layer
|
|
"What the technical customers want is control over their compute, their models, their data stack, and their alpha. They want to know they own the means of production, and it's not being transferred to someone else....this is the voice of American business that is being channelled through me.” - Palantir CEO Alex Karp on CNBC
Alex Karp was interviewed on CNBC recently, and went semi-viral for firing a number of broadsides at closed-source AI labs in general and Anthropic in particular. The background context to the interview was the announcement that Palantir is partnering with NVIDIA to build applications on their fully open source Nemotron model. His arguments for open over closed source centre on three main issues: data protection, customer sovereignty over AI models, and, separately, a broader point about the ROI and value proposition from an AI buyer's perspective. Palantir later put out a more detailed piece on the data protection/proprietary knowledge side, and Microsoft's CEO Satya Nadella has been making similar points, most recently in a blog post titled " The Reverse Information Paradox".
These are not entirely disinterested interventions - both Microsoft and Palantir benefit from an architecture that commoditises the model layer and elevates the operational software surrounding it. On the flipside, it is in the interest of frontier labs to convince entreprises that frontier intelligence remains scarce, difficult to reproduce and sufficiently differentiated that they should optimise around access to the best model rather than insist upon ownership.
As always there is lots of nuance to these debates, which we will try to unpack here, as well as offering our take on the tension between open and closed source models, the rising importance of the orchestration layer, and implications for the AI infrastructure build out.
|
|
|
Palantir points to their classified military work with government or highly regulated industries as examples of domains where data confidentiality is critically important; in fact, asks Karp, how can any entreprise justify letting third-party language models ingest their 'alpha', and potentially incorporate proprietary IP into future LLM training runs?
This is not a new debate - AWS just turned twenty years old, and the whole of the 2010s were spent discussing 'Big Data'. Flexera estimate 53% of Entreprise and SMB data and 56% of workloads are stored and run in the public cloud today, and organisations are becoming increasingly comfortable with migrating across ever more sensitive data. For those that aren't, there are numerous hybrid cloud options, and, in extremis, hyperscalers can build bespoke facilities, such as AWS' GovCloud (US-West) facility which launched back in 2011. Since then AWS and Azure have extended similar services to the US intelligence agencies and defence contractors.
The frontier AI labs are already following this well-trodden path - after all, DeepMind is 100% owned by Alphabet, OpenAI is 27% owned by Microsoft (and partnered with Oracle on Stargate) and Anthropic is 30% owned by Alphabet and Amazon. All of these big tech companies have decades of experience handling sensitive data for governments and highly-regulated industries. Azure OpenAI Service and AWS Bedrock (via PrivateLink) provide private servers or endpoints which allow entreprises to run OpenAI and Anthropic's latest models on their own virtual private clouds. The data never traverses into the public internet, and the hyperscalers contractually guarantee that customer prompts and outputs are never used to train the labs' foundation models. For the vast majority of entreprises and SMBs, contractual terms of service and SLAs are perfectly strong enough to protect their data. Even basic consumer subscription accounts have easy toggles to opt out of data collection.
|
|
|
Data protection can range in strength from self-owned, air-gapped hardware to weaker contractual arrangements in the public cloud
|
|
|
|
Source: Palantir
|
|
|
That doesn't mean AI labs aren't training models on user data - the devil is in the detail with these contracts. Maybe uploaded files are protected, but what about user prompts? Today's models have 800,000 word context windows - millions of long, well crafted prompts and subsequent back and forth conversations may provide valuable insight into private corporate workflows and the reasoning process behind them. Satya Nadella talks about "intelligence exhaust" - knowledge that leaks imperceptibly through the tiny, everyday corrections and evaluations employees give the AI when it gets something wrong. These metadata can, and probably will, find their way into the training data corpora of future models.
This matters a lot to some companies and government agencies. For closed source model use, the solution is for these organisations to use their clout to negotiate bespoke zero data retention agreements (ZDRs) directly with the AI labs. But how significant a proportion of AI use are these organisations? Will consumers demand such protections? Probably not - as we have seen with social media, consumers are quite happy to volunteer their data if it keeps the cost of the service low. Similarly, prosumers using chatbots for work, or even SMBs, are likely to go with the default options, perhaps toggling on the weaker form of data protection in their settings when they first set up their account. Only 18% of the US workforce works for an S&P 500 company - this leaves 130 million workers, counting among them millions of professionals with similar expertise and knowledge to those Palantir is seeking to protect from data harvesting. Not only that, hundreds of thousands more are being enlisted by companies like Mercor with the express purpose of creating clean datasets of every professional workflow imaginable to sell to frontier labs.
Against this backdrop, how much alpha can actually be protected by ZDR policies? Is the tribal knowledge inside a magic circle law firm or tier 1 investment bank so different to smaller players in their industry? Many of the employees of smaller firms began their careers at larger institutions. Of course, valuable proprietary data exist - IP, trade secrets, pharmaceutical R&D - but this isn't necessarily what AI labs are trying to harvest. Rather, they want to capture the longer time horizon reasoning traces of lawyers, financial analysts, engineers and scientists to increase the intelligence of their next model. Each expansion of the jagged frontier of intelligence unlocks new, economically valuable use cases.
|
|
|
AI labs aren't seeking to extract proprietary data to crystallise information in model weights, they are trying to improve generally transferable reasoning in models to smooth and expand the jagged frontier of intelligence
|
|
|
Customer Sovereignty over AI models
|
|
|
Palantir advocates for entreprises to use open source models as a means to avoid upstreaming value to the closed source labs. This addresses the data protection issue, but also provides a defence against vendor lock-in and can also reduce costs.
|
|
|
An entreprise that controls their own model not only protects their IP and knowhow as it stands today, but can implement a flywheel of improvements that compounds value over time
|
|
|
|
Source: Palantir
|
|
It seems like a great idea, but is not as easy as it sounds. Even Microsoft - the largest entreprise software company in the world, the #2 hyperscale cloud provider, and with full, unrestricted rights to OpenAI's IP - has struggled to integrate AI into their organisation in the way Palantir is suggesting. The next largest software firms, the likes of Salesforce et al, have barely progressed beyond basic implementations of chatbots, despite making AI their #1 priority for the last 2-3 years. Even the straightforward-sounding task of fine-tuning open source base models on proprietary data is much harder than one might imagine, and beyond the capabilities of most organisations ( as explained by Andrej Karpathy here).
That isn't to say it's impossible - Thinking Machines published research in partnership with Bridgewater that showed fine-tuning a model on proprietary, expert-labelled financial data could improve accuracy by +650bps at 13.8x lower cost on some financial tasks (namely judging whether financial articles, central-bank documents and research reports were relevant; distinguishing one-off analysis from recurring boilerplate; and identifying where useful content ended in documents or emails). While interesting and possibly useful, the report lacks detail on whether model performance more broadly was impacted by the fine-tuning (as it often can be). Furthermore, the customised model is frozen forever - fine if it gets the job done and requires no improvement, but if Bridgewater want to swap in a newer/smarter/cheaper model down the line, the fine tuning process has to start all over again.
As we write this piece, Thinking Machines have released their own open source model, which is roughly the size of the current crop of leading Chinese ones (975B parameters/41B active MoE) and looks to be a new leader in US open source, edging out Nemotron 3 Ultra. Their angle is to build tools that make it easier for customers to fine tune and customise their models, so perhaps this will help lower the barriers to entreprise embracing Palantir's vision.
|
|
|
Thinking Machines and Bridgewater demonstrated higher performance and lower costs from fine tuning an open source model on a narrow subset of financial tasks
|
|
|
Open Source vs. Closed Source
|
|
|
It's worth refreshing on the current state of play in model capabilities, given the prevailing narrative in the markets that Chinese open-source models are nipping at the heels of the frontier labs, offering similar performance at far lower cost.
At the start of the year, Epoch AI estimated open source models were lagging closed source by four months. The proviso here is that open-source models often score much worse on private benchmarks, suggesting they may be trained to score well on public leader boards, but have more 'jagged' intelligence when broadly applied to useful tasks. Open source tend to be behind in multi-modal reasoning, lacking 'big model smell' - a subtle quality of rounded, general intelligence that only becomes apparent after prolonged use. Against this measure, some put the gap at more like six months, and possibly widening. Like closed source frontier models, coding capability has improved the fastest, due to its amenability to reinforcement learning in the training process.
|
|
|
Epoch AI estimates open source models are lagging closed source by around 4 months
|
|
|
Large, closed source models still dominate the frontier, especially the most recent releases from Anthropic (Mythos/Fable) and OpenAI (GPT-5.6 Sol). But Chinese open source have been fast-following, and it is generally agreed that part of this has been through the distillation of larger US frontier models. Anthropic have been vocal on this topic, and are lobbying for an intervention from Washington. Most recently, they wrote a letter to the Senate, alleging: "Between April 22 and June 5, 2026, this campaign generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts, in violation of our terms of service and access restrictions. Alibaba’s campaign targeted some of Claude’s most valuable capabilities, such as agentic reasoning, software engineering, and long-horizon tasks. Congress should advance measures that facilitate threat information sharing between US AI labs, close loopholes allowing PRC AI labs to access advanced US chips, and penalize PRC labs responsible for distillation attacks". There are rumours of an executive order on restricting Chinese open source models due to be announced by the White House shortly.
|
|
|
It's not just Alibaba. Anthropic have accused nearly all of the leading Chinese AI labs of distilling their models over the last few months.
|
|
|
|
Source: X
|
|
|
The results of distillation can be easily seen within model families - OpenAI's GPT 5.6 Luna, Google's Gemini 3.5 Flash and probably Meta's Muse Spark 1.1 were all trained by larger 'teacher' models in their family. And their aggregate benchmark scores are suspiciously close to the leading Chinese open source models of a similar size. We are about to see a new crop of ~3T parameter Chinese models released (e.g. Kimi v3), which may be competitive with GPT 5.6 Sol; while Fable/Sol have only just been released broadly, it should be noted that they are five months old, and their successors are already being tested internally.
|
|
|
Artificial Analysis aggregate benchmark scores for select closed-source models (blue) and open source models (green)
|
|
|
|
Source: Artificial Analysis; Green Ash Partners. Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
|
|
|
Chinese models will never be able to surpass the frontier through distillation alone, but there are genuinely strong AI research teams in the leading Chinese labs so we expect the ferocious competition in the AI race between the US and China to continue. The US may yet extend their lead due to their advantage in compute infrastructure, and we do think the US government may step in to restrict the use of Chinese models in US companies on security grounds (much like they did with Huawei in mobile networks). But there are US open source models too - Palantir's recent comments on "owning the means of production" came in the same week as an announced partnership with NVIDIA to use and develop their open source Nemotron model family.
Data privacy aside, another reason to use open source models is to save on cost. AI labs are reportedly selling tokens at ~80% gross margins - similar to the SaaS software gold standard - and then there is the cloud margin stacked on top. Why not self host open source models to save on token costs? A few companies may try this, just as some industries, such as banks, run substantial amounts of their own on-prem datacentre infrastructure, however it requires considerable resources both in terms of capital and technical staff.
As it happens, model releases from OpenAI, Meta and SpaceXAI in the last few months have reclaimed the Pareto frontier in performance versus cost from the Chinese open source labs. It is notable that Kimi V3, likely similar in size to GPT-5.6 Sol, is also very similar in cost ($0.94 vs. $1.04 per task in Artificial Analysis benchmarking). The model with the cheapest token per unit of intelligence will shift with each new release, but, closed or open source, as a general rule the cost per token for a given level of intelligence declines by about 60x per year - any workflow that is too expensive today due to the cost of running a frontier model will be affordable, or even cheap, by next year.
|
|
|
GPT 5.6 Luna, Muse Spark 1.1 and Grok 4.5 - all released in the last month - have edged out Chinese open source models on the Pareto frontier of performance vs. cost
|
|
|
|
Source: Artificial Analysis. Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR
|
|
We have touched on the idea of harnesses in previous writings ( On the Horizon #8: The Age of Agents; and, Applied AI: The Great Unhobbling) - examples include Claude Code or Codex. Harnesses decide how work gets split, which subagents spawn, what tools each one gets, how output is verified, which model handles which step, how work is isolated, and when the job is done. The orchestration layer coordinates multiple harnessed AI agents across an organisation. In essence, the models are becoming cognitive engines inside a much larger operating system.
Palantir's vision is for entreprise to own this part of the stack, both in order to optimise compute costs, as well as to create a kind of positive improvement loop based on an organisations proprietary workflows.
|
|
|
Palantir exhorts entreprises to own their harness and create modular processes that can readily slot models in and out
|
|
|
|
Source: Palantir
|
|
|
As with the customer sovereignty over AI models section, this is easier said than done, and even in the software sector, very few companies have demonstrated the ability to create model harnesses that are competitive with Claude Code or Codex. Introducing models from multiple providers introduces further complexity, and transferring context management/long term memory/skills and plug-ins across harnesses introduces bugs and performance issues that entreprises will not want to risk inviting in to their critical workflows.
Currently the main beneficiaries of harness data are those with the largest numbers of users - the AI labs. A harness turns vague, multi-step agent workflows into structured data that can be logged and graded, allowing labs to hill-climb effectively, and the frontier of post-training is now RL directly on agentic trajectories inside these scaffolds; scaffolds are therefore temporary - they are compensations for current model weaknesses, absorbed into the next model. This is very valuable, as evidenced by SpaceX's recent acquisition of Cursor for $60 billion, and their return to (near) the frontier shortly afterwards.
|
|
|
Codex has grown from one million users to nine million in five months; Claude Code is reportedly approaching five million users
|
|
|
|
Source: Green Ash Partners
|
|
|
Entreprises can own and control the router - this sits above the harness, and decides which model is best suited to handle a task. It is primarily an optimisation tool, and, while it can save significant costs, it is a less valuable slice of the orchestration layer than the harness, given the latter's amenability to continuous capability improvement through data ingestion. Routers can also materially degrade performance - GPT-5 launched with a router in August 2024, with the reasonable goal of preserving scarce compute by serving a small model for every day queries. But because the router operated as an invisible black box users had no control over - or insight into - which "brain" was handling their request. This led to a broad perception that GPT-5 was a bad model and that AI progress was slowing, which persisted well after the router was removed (two months later, reasoning models kicked off a step-change leap in capability). We have seen some similarly poor implementations of routers in AI wrapper start ups aimed at finance professionals. These companies are incentivised to send queries to smaller models to lower their inference costs and protect margins, and in doing so give users a misleadingly poor impression of what frontier models can actually do.
|
|
|
Model routing and usage optimisations at Coinbase supported continued overall token usage growth at ~-40% lower cost
|
|
|
|
Source: Coinbase
|
|
|
The Implications for the AI Infrastructure Build Out
|
|
Entreprise scrutiny on AI spend has been framed as a negative demand signal for token consumption, and/or a lack of confidence about the ROI of AI adoption, but it is far too early in the adoption cycle to start tracking these kinds of metrics. As we wrote in our newsletter last month, data from OpenAI shows the relationship between the uptake of agentic workflows and token consumption (their internal consumption has increased 10-50x since the start of the year, depending on the department). Outside of AI-native companies agentic AI is in the very earliest part of the adoption curve.
|
|
|
Despite all of the talk of tokenmaxxing AI agents in entreprises, AI agent use is still in its very earliest innings
|
|
|
Yes, there have been stories in the press about this or that company capping usage or scrutinising costs - like Uber capping software engineers at a $1,500/month token budget or Tesla instituting a $200/week cap - but per Ramp data, the median monthly AI spend per employee is just $11 per month (and 40% of their customers are tech companies).
|
|
|
The Top 1% spend ~12x more than the Top 10% and 650x more than the median
|
|
|
|
Source: Ramp
|
|
|
There is another bearish framing around the close source/open source debate, which is a sort of lingering after effect of the "DeepSeek Moment" which was largely a story of efficiency and cost optimisation in model architecture. Algorithmic improvements are discovered and adopted by all of the labs all of the time, but Chinese labs today are very much on the same scaling trajectory as US ones.
|
|
|
DeepSeek, MoonShot (Kimi V3) and MiniMax are all stepping up to larger model sizes
|
|
|
This is bullish for AI compute and the AI infrastructure theme. Whatever combination of models companies end up deploying, there is no getting around the compute intensity of long-running, context heavy agentic workloads. NVIDIA will produce in the order of 6 million GPUs this year, or ~10GW of compute. This would be enough to serve the entire ~1 billion-strong LLM chatbot user base, if everyone used a DeepSeek v4-sized model for short context, ad hoc queries (~1T parameter, sparse MoE). But increase the size of the model to a denser, 10T parameter Mythos-class, and make full use of a 1 million token context window, and this same 10GW could only serve ~34k agents running concurrently.
|
|
|
One year of Blackwell production (~6 million GPUs, ~10GW) supports anywhere from ~1.3 billion chatbots to just 34k frontier agents, depending on model and context
|
|
|
|
Source: Green Ash Partners estimates; Epoch AI calibrated inference model; Artificial Analysis AA-AgentPerf; NVIDIA. Assumes ~83,000 NVL72-equivalent racks, 20 tokens/s per agent SLO, INT4 weights, FP8 KV cache, MLA-style attention for sparse classes. 1M-context figures extrapolated beyond published benchmark range; denser-MoE 1M figures are generous. Illustrative.
|
|
|
Of course, algorithms will improve efficiency, models will be distilled into smaller formats (consensually or otherwise) and hardware will improve (NVIDIA's Blackwells can run 20x more concurrent agents than Hopper at 20 token/s). But so too will models continue to scale, with NVIDIA's subsequent Rubin and Feynman systems being designed with 100 trillion parameter models in mind.
The longer agents can run, and the wider the parallelisation of multi-agent systems, the greater the potential for AI to take on high-value economic tasks, and so the drive to train larger and larger models will continue as long as capital continues to flow and scaling laws continue to hold.
|
|
|
NVIDIA GB300 NVL72 supports 20x more agents per MW than NVIDIA H200 at 20 and 60 tokens/s per agent
|
|
|
Forecasting a Rapidly Moving Frontier
|
|
|
The history of the LLM debate is already littered with supposedly terminal limitations: Models produce nothing but slop; hallucinations make them unusable; they can't remember, reason or plan; next-token prediction can never support genuine intelligence; agents fail as soon as a task extends beyond a few minutes. Each critique identified something real about models at the time, but made the same forecasting error - treating a rapidly moving system as static.
Weaknesses that looked architectural were mitigated by scale, reinforcement learning, inference-time compute, retrieval, memory, tools, verification and orchestration - and the most effective scaffolds are increasingly being folded back into the next generation of models.
None of this tells us precisely when artificial intelligence will transform a given industry, or replace a given job, nor which model architecture will win or how the economic value will be divided. But we suggest that extrapolating today’s limitations in a straight line is an unreliable way to forecast what comes next. The one assumption we can rely on is continued rapid improvement. And, that as models become larger, agents run for longer, contexts deepen and millions of workflows are parallelised, there will be an enormous continuing requirement for compute.
|
|
|
AI capability improvements are not just improving, but the pace is accelerating. It is human nature to extrapolate exponential trends in a linear fashion, assuming models today are as good as they are going to get
|
|
|
|
Source: METR, the discourse, oneusefulthing.org
|
|
|
|
|
This communication is issued and approved by Green Ash Partners Investment Management Ltd ("Green Ash"), which is authorised and regulated by the Financial Conduct Authority (FRN 1015503). Registered in England and Wales, company number 14963372. Registered office: 11 Albemarle Street, London, W1S 4HH.
|
|
|
|
|
|
|
|