The biggest takeaway this month is the re-release of Fable 5 and the new release of ChatGPT 5.6 Sol. These frontier models introduce a new paradox in AI performance: they are so capable that they almost make you take your hands off the wheel while you're still driving. The capabilities are incredible, but the models still require strict guidance to get to the outcome you specifically want - versus the outcome they discover on their own.
That distinction - the outcome you want versus the outcome the model finds - is the thread running through everything this month, from how the day-to-day work has changed, to the open source releases out of China, to why I think both public and private markets are misreading what comes next.
The 80% Era Is Over
With prior models, you had to go through several rounds of iteration to get an output you truly wanted. They have been super capable for the past 8 months or so, but the work followed a rhythm: build a plan, get the model to understand your context, turn the plan into an execution exercise, iterate on it together, and then set it to work. Whether you were building in code or working on a PowerPoint, it was a multi-turn affair. The first output would be really good - but definitely not complete and production-ready as a one-shot exercise, meaning one prompt in, finished product out, no iteration.
This is why a lot of people would tell me, "We tried using Claude in Excel and it produced about an 80% solution - not good enough to use for the real financial models." My comment back was always the same: they just hadn't taken the time to realize that with 2-3 more turns of prompts they'd land on their target output - and then they could turn that multi-turn process into a skill they could run repeatedly as a one-shot tool in their toolkit. The iteration wasn't a flaw. It was the setup cost for automation.
With these new models, the difference is that the one shot produces the 100% solution. It just might not be the solution you expected.
The models now have such a high degree of autonomy and self-direction that they can run for 12+ hours on a task once you tell them what you want produced. If I ask for an app to track my wardrobe, a one-shot prompt will build it for me - as long as I'm ok with the model making its own choices on design, user experience, and everything else. It will be an amazing app. It will also, potentially, be a different app than the one I pictured, unless I gave it very explicit guidance up front.
A 180 IQ Intern Without the Domain Judgment
So as I've used these models for work outputs, and as they've proven their capabilities, I've found myself giving them harder problems to work on by themselves. What I found is that this is an incredible superpower - so long as you've taken the time to define the specific output you need, with clear measures of success and rules of the road for execution.
As I explained to one person, it's like a 180 IQ intern joining your firm. The intelligence is unbelievable. But the intern hasn't come up the curve on domain knowledge, or on the proprietary judgment of what high quality - what "taste" - looks like in your business. It will produce incredibly robust reports on a company or a person, and still miss the critical signal you actually needed it to focus on, delivered in 1 page rather than 10 deeply researched pages.
So we adjust and we learn.
And we find the growing value of our own data, context, and judgment - the raw material that steers these models to work for us effectively. This is exactly the shared memory and context layer I wrote about last month: the models keep getting smarter and cheaper, while the thing that makes them useful to you compounds privately, on your side of the table.
Open Source Doesn't Say No
If that was the initial story of the month, the bigger one followed: the release of the latest set of open source Chinese models from Moonshot AI and Alibaba, called Kimi K3 and Qwen 3.8. These models essentially reproduce the performance of Fable and 5.6 Sol at a fraction of the cost, and they can be run on your own infrastructure - without your work feeding back to Anthropic and OpenAI, and with the cost of the frontier radically reduced.
The trend itself is nothing new. Chinese labs have been a couple of months behind the US leaders all along. Back in May I wrote that we would probably have an open source model approaching Mythos-level capability within a few months. It took about two. But the sheer performance we've now reached with these models has made the story much louder.
Remember: Mythos and Fable were held up by the US government because they were afraid the models were too powerful to be released without oversight and safeguards in place - and GPT 5.6 went through the same gate. Through that process, the models have even been restrained a bit, refusing to do some work they think is dangerous, even when it isn't. It also meant the 2-3 month lead the American labs had was eaten up by this regulatory process.
Chinese models have no such guardrails. And they don't say no. And since they're open source, they can be fine-tuned against any biases they may come with out of the box.
It is a huge turning point.
These latest frontier models are pushing capability up while open source pushes the cost of that capability down - which means we will find way more work we can delegate, and way more places to put these models and agents to work.
Cheaper Tokens Don't Mean Less Demand (Jevons Paradox)
On this last point, I find myself really puzzled by the investment landscape, on both the public and private side.
All of these dynamics - models that can run longer and more autonomously, on more tasks, at a cheaper cost of inference (the cost of actually running the models, as opposed to training them) - point in one direction: demand is going to explode. The narrative around companies checking their AI spend and reining it in is not a dynamic where they're using AI less. When you swap in a cheaper open source model, you're still consuming as many - if not more - of the lower-cost tokens that serve as the real measure of compute demand.
So companies will be using AI even more; they're just spreading the inference tokens across a wider array of providers. And all of that demand still feeds an insatiable buildout of data centers, from the neoclouds - the specialized operators built purely to rent out AI compute - to the hyperscalers alike, the same buildout I detailed in May. Critically, this demand doesn't depend on any individual model winning. We are seeing more and more models, better and better and cheaper, constantly - and they all need a place to run.
Memory Is the New Nvidia
All of this points to more chips - and critically now, more memory demand than the world can produce.

Micron fell 35% recently despite posting massive quarterly numbers
Which is why the idea circulating in public markets that memory stocks are becoming overvalued confuses me. When it comes to memory, no one can get enough of it, and Apple proved the point this month by raising prices to pass these costs on to consumers. It feels like Nvidia in 2023, when the stock seemed expensive because it was so untethered from prior demand - and it's now multiples more valuable. Memory feels like the same deal.
Now I’m not a trader or public markets investor, so take all of this with a grain of salt. There may be some bigger systematic risk in the markets if the dynamics I’m laying out here end up crippling Anthropic or OpenAI who’ve committed trillions to capex plans. But I generally think even in an extreme scenario when one of these companies fails somehow, Google, Meta, Amazon and everyone else will quickly scoop up the commitments they leave behind.
The Shrinking Development Moat
For the startup world, though, all of these advances feel like a huge risk.
Now I’ll admit this theory isn’t new and is the whole reason for the SaaSpocalypse narrtive. But as one executive told me when I showed him a demo I had created recently: "Doesn't this mean we could just kill our current CRM?" The answer was probably yes. And the same goes for many SaaS apps - though not necessarily the ones people have been worried about. Many many SaaS apps have little to no risk of disruption here. In fact they may become much more valuable.
The vulnerable companies are the ones without deeply entrenched enterprise adoption, the ones that don't sit as a critical system of record - the software that holds the data a business cannot operate without - and so can be easily yanked out. That's consistent with what I argued back in March: deeply embedded systems of record are resilient; it's the software around them that reprices.
The bigger threat is to the huge number of startups that have seen exponential growth, where all of a sudden people look at the product and say, "Hey, I can build that too" - and open source alternatives to everything start popping up. The moat to development is zero.
What does have value is optimizing the user's stack: better routing, better hosting, better integration, better memory and context. Those things are sticky. Anything built on top of them feels like a potential house of cards. That's why I'm so interested in the agentic stack I've been writing about - and why I'm working with businesses in the services world that benefit from the improvement in models coupled with the improvement in tooling, changing cost structures in verticals where you still have a long runway to become the category leader versus the more tech-forward crowd.
Thanks as always for reading,
-Jake
Jake Dwyer
Founder & Managing Partner
Factor Capital
