Open models, open questions
Is there an overlooked business opportunity in open models?
A short while ago, I was talking to an AI engineer working on several million-dollar AI deployments. He mentioned there’s an overlooked opportunity in pointing business problems at open-weight models, but nobody wants to talk about them, preoccupied with loops and agents. Myself being very much preoccupied with loops and agents at the time, I let the comment pass. But I was reminded of it two weeks ago when my twitter timeline exploded with open models.
There’s the Jensen Huang’s first-ever tweet to a not-a-protest, 450-RSVPs mini-march through San Francisco, and a plot twist of the HuggingFace hack: commercial closed models refused to assist in investigation, their guardrails careful as to not be foiled by a smart hacker, pretending they’re a victim. HuggingFace’s self-hosted open model, came to the rescue.
Most of the defense of open models hangs on the principles: open source, innovation, free markets and all. There are more business-model based incentives behind those principles, but if you’re on the hook for making AI work in your company, ultimately what matters for you is: what do I gain from open weights?
Nobody gets fired for buying IBM relying on closed models
Open models do seem to be an overlooked opportunity - most roadmaps focus on self-improving agents with whichever API is performing the best in the result-vs-cost matrix. You probably have a rough idea of the benefits of open models: they can be cost effective, their privacy and data residency, lower dependence on the whims of external platforms, and some flexibility on restrictions on what work the LLM is allowed to do.
But a bigger reason I now see them as overlooked is that they offer a tempting benefit for companies who need something a little more custom than the out-of-the-box API.
Here is a potential failure mode: generic post-training skews towards producing responses that are maximally acceptable to an enormous range of users. This is an impressive feat! The same model that’s designing a marketing strategy for a Fortune 500 company can produce a similarly acceptable response for a vibecoder with an app.
Now, if you are running a business with an idiosyncratic set of needs, this feat might be the exact issue you’re stumbling over in deploying AI: what if your aerospace maintenance instruction document needs to fit some very specific writing standards? Or if every security alert has to be matched to some internal taxonomy and escalation?
Closed models can and do handle this work, with mixed results. Whether open-weight models are the right answer is a decision somewhere in the neighborhood of feasibility and risk appetite, because fine-tuning a model is 💸 expensive 💸, difficult, and clear examples of ROI are scarce. The costs include: a team of experts that can do it and maintain it over time; hundreds or thousands of examples that someone with expertise (💸) manually curates, format, evaluate, and then maintain. In return for the investment: a yet-unvalidated hypothesis that fine-tuning might work out as an advantage. It might be the anti-IBM character in the “nobody got fired for buying IBM” allegory: when everyone’s calling APIs, the bet of customizing a model seems unjustifiable to make.
There’s also a unique risk of obsolescence to custom models, and BloombergGPT come to mind. A bit of an imperfect example, since it was trained as a new model, but the risk still stands: Bloomberg trained their own model on loads of proprietary data in 2023, with the goal of offering an edge in financial tasks. It outperformed similarly-sized models it was compared to. Unfortunately, frontier models moved quickly, and BloombergGPT did not. Soon after BloombergGPT came out, an external paper showed that GPT-4 comfortably outperformed it on public benchmarks1. So the ROI on whatever investment you make into customizing a model will also compete with the yet-unknown future capabilities of frontier models. Which you’ll be able to get for the cost of an API call.
I napkin-mathed an interactive prototype on open model economics for a midmarket business - pls help me shape the future topics by answering this short survey and I’ll send you the link to the model.
Long-term feasibility is an open question and - perhaps? - the next breakthrough topic. I am curious about use cases where models customized by run-of-the-mill businesses turned out to be ROI-positive, especially in the long run. And whether in 6-12 months the topic of LinkedIn posts will shift from self-improving agents to customizing open weights as the latest advantage in AI-native. Mira Murati’s Thinking Machines seems to be making a bet on exactly that.
Meanwhile, for the vast majority of use cases, a good harness or a Forward Deployed Engineer is often all that’s needed.
From the chatter
There might be internal tasks that BloombergGPT was still better at, bringing some ROI.






