<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Artificial Confidence]]></title><description><![CDATA[What changed in AI this week, adjusted for spin.]]></description><link>https://artificialconfidence.com</link><image><url>https://substackcdn.com/image/fetch/$s_!LrGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F359fc899-ddd0-4b25-950b-95adb8a34ae1_1024x1024.png</url><title>Artificial Confidence</title><link>https://artificialconfidence.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 14 Aug 2026 12:16:55 GMT</lastBuildDate><atom:link href="https://artificialconfidence.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Corey Quinn]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[artificialconfidence@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[artificialconfidence@substack.com]]></itunes:email><itunes:name><![CDATA[Corey Quinn]]></itunes:name></itunes:owner><itunes:author><![CDATA[Corey Quinn]]></itunes:author><googleplay:owner><![CDATA[artificialconfidence@substack.com]]></googleplay:owner><googleplay:email><![CDATA[artificialconfidence@substack.com]]></googleplay:email><googleplay:author><![CDATA[Corey Quinn]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Reading the tea leaves: 2026 predictions]]></title><description><![CDATA[In which I get to make some predictions about our space.]]></description><link>https://artificialconfidence.com/p/reading-the-tea-leaves-2026-predictions</link><guid isPermaLink="false">https://artificialconfidence.com/p/reading-the-tea-leaves-2026-predictions</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Wed, 29 Jul 2026 01:20:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LrGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F359fc899-ddd0-4b25-950b-95adb8a34ae1_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Something I never got to do in my past life as an AWS commentator is "predictions." It turns out that when you have access to confidential information, you can't very well make predictions that touch on those things without a serious lapse of ethics. Thus, any AWS prediction I make fell into two horrible failure modes: either I get it wrong, which is the least bad option, or else I get it right and everyone thinks I knew but somehow leaked other people's information, in which case I don't get to have a business anymore.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:null,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free; I won&#8217;t leak your information either.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p>I don't have that problem in the AI space, because nobody's giving me their roadmaps. I'm not entirely sure I'd trust any of them, given how rapidly this space is evolving. And so, given that today's my birthday, I'm going to indulge myself by predicting things about the world of AI that may or may not come to pass, driven by overall trends.</p><h2>The marginal cost of inference falls to zero</h2><p>With open models rivaling frontier models, there&#8217;s a strong possibility that the cost of inference approaches a tiny premium over the cost to run its infrastructure. I'm not suggesting that you're going to run Kimi K3 on your laptop anytime soon, but folks like Baseten (who you should be paying attention to) and others are absolutely going to provide it to you on a cost-plus basis that makes things like Opus (let alone Fable) 5 pricing look like someone's retirement plan.</p><p>I'm also predicting that running a frontier lab whose primary business is selling inference against SOTA models will start to be shaped like a business whose margins are less "SaaS" shaped and more "suspiciously close to an airline's." Hence, I see the current scrabble as OpenAI and Anthropic both looking to love up the stack, and in so doing they start to look more like the AWS of yesteryear&#8212;backstabbing of "partners" and all.</p><h2>Open harnesses won&#8217;t suck</h2><p>Right now, both Codex and Claude Desktop alike are very similar, which is a polite way to say that they're both not great products, which is me desperately trying to upset friends at both companies by referring to their products as something scatological. I've lost count of the conversations I've had here in San Francisco with folks who're building their own agent harnesses. I think that this phase is going to pass (good god, it's not like we want to roll our own text editors either), and we're going to see a future iteration where the most common consumption pattern is an open harness or two around which the industry congeals like grease on a stove. Transparency around system prompts, token consumption, data exfiltration concerns, and more all point to this being increasingly important to the enterprise, whereas the developers want outcomes, not a weekend project just to wrangle the tools.</p><h2>The fundamental nature of computers changes</h2><p>The big shift that AI drives that I haven't heard articulated much is that for the first time since the dawn of computing, computers do what you <em>mean</em> instead of what you literally <em>say</em>. (Also somehow they are bad at math now.) That one shift reframes the entirety of computing from "know the magic incantation to make the computer do the specific thing you want" to "barely articulate the outcome you're after" as a viable path to an outcome.</p><p>We&#8217;re already seeing that developers aren&#8217;t getting tripped up on the syntactic nonsense they used to; in time, knowing how to write code by hand is going to be about as germane to software engineering as knowing assembly is today. Instead, the valuable marketable skill will be in how to coax the models into doing your bidding. That may not sound like a job (it is absolutely your boss&#8217;s job), and yet you&#8217;d also not think &#8220;knowing how to coax information out of a search engine&#8221; would have been a decades-long durable skill either.</p><h2>Neoclouds are on a ticking clock</h2><p>The reason the neoclouds are in business today is almost entirely due to supply constraints. There are virtually no enterprises who are <em>pleased</em> to be doing business with them; this is where they find themselves when they attempt to get GPU capacity from their usual hyperscale cloud providers, get told "get in line," and have business needs that won't wait on supply line physics. At some point in the future, supply constraints will ease and you'll be able to mostly get all the GPU you care to, from any provider you care to, much the way CPU bound instances work today. When that day comes, the neoclouds are going to have to have built a differentiated offering beyond pure GPU unless they want to follow in the wake of the swath of VPS providers that lost their market with the rise of cloud.</p><h2>Model routing becomes a primitive</h2><p>The current state of every piece of software chooses where to run its inference seems less sustainable to me than other paths. I can see a world where &#8220;inference&#8221; becomes an operating system interface that applications use, and based upon a variety of factors (how complicated is the task, is there currently an internet connection, has the user expressed a lack of caring about money) routes it accordingly. We&#8217;re already seeing model routers doing &#8220;smart routing;&#8221; bringing that decision tree closer to home yields many benefits.</p><h2>SaaS will continue to thrive</h2><p>The open source world learned previously that software is free as in puppy. We&#8217;re now learning that while vibe coding is fun, vibe maintaining sucks, and vibe maintaining someone else&#8217;s slop is just the absolute worst.</p><p>There&#8217;s strong business value to having a team of people care deeply about the problem domain, have the expertise to recognize when models are going off the rails, and can imagine your future needs in such a way that &#8220;rewrite the app when requirements change&#8221; isn&#8217;t your only path.</p><h2>You will still have a place</h2><p>Despite all the hoopla about automating folks&#8217; jobs away, evidence of this happening at scale remains thin. Block <a href="https://www.cnn.com/2026/02/26/business/block-layoffs-ai-jack-dorsey">claimed their 40% layoff was due to AI</a>, but I&#8217;m a skeptic; I note they didn&#8217;t do big layoffs when all of their technical peers were doing them 1-2 years earlier, and &#8220;we&#8217;re the first of a new thing&#8221; absolutely plays better than &#8220;we&#8217;re staggering across the finish line hours after the race is over and the organizers have gone home.&#8221;</p><p>The shape of some jobs may change; the tools we use as technologists certainly will. But the two things that can&#8217;t be industrialized are wisdom and judgement, and it&#8217;s becoming increasingly clear that LLMs aren&#8217;t on a path to changing that. We&#8217;re in for some interesting times in the next few years, but we&#8217;ll still be here when the turmoil settles back down.</p>]]></content:encoded></item><item><title><![CDATA[Frontier lab economics have bifurcated: tax the exit, lobby against the free]]></title><description><![CDATA[Distillation is cost optimization at home and theft abroad.]]></description><link>https://artificialconfidence.com/p/frontier-lab-economics-have-bifurcated</link><guid isPermaLink="false">https://artificialconfidence.com/p/frontier-lab-economics-have-bifurcated</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Tue, 21 Jul 2026 23:07:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gTlw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Tuesday morning saw <a href="https://www.cnbc.com/2026/07/21/bessent-china-ai-sanctions.html">US Treasury Secretary Scott &#8220;Chucklepony&#8221; Bessent on Fox Business</a> saying that the administration will investigate whether Chinese models were distilled from American ones; he left sanctions on the table for &#8220;theft.&#8221; The obvious answer half the internet is responding with is aligned with &#8220;you&#8217;ve attempted to steal what I have rightfully stolen.&#8221;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free so I can attempt to steal your heart.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p>Earlier this month, Kimi K3 shipped, which is the largest open weight release to date. This was met by the model labs and other ill-informed blowhards calling for regulatory action:</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/jimcramer/status/2079509493805642189&quot;,&quot;full_text&quot;:&quot;We won't let the Chinese use our Nvidia chips, even as that would have made them dependent upon us. But we are willing to give them all our corporate data? Really? I know Fintech. This is Finsuicide&quot;,&quot;username&quot;:&quot;jimcramer&quot;,&quot;name&quot;:&quot;Jim Cramer&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1461426655046606860/PzlSk4fZ_normal.jpg&quot;,&quot;date&quot;:&quot;2026-07-21T10:11:13.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:297,&quot;retweet_count&quot;:40,&quot;like_count&quot;:614,&quot;impression_count&quot;:190534,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p>This... is cute, given that these open weight models are being hosted by US-based providers, not Chinese providers or the labs themselves. There&#8217;s no &#8220;sending your corporate data to China&#8221; aspect here. But there is an economic one: my thesis is that this is a pricing story that makes more sense once you go back to a leaked contract from three weeks ago.</p><h2>The Amazon contract</h2><p>The Information reported that as of next year, Anthropic has reworked its Amazon contract to <a href="https://finance.biggo.com/news/2b5b2756-ef58-49c7-87a6-2927875b55b9">reprice Amazon from compute-hours to tokens</a>; this is the same structure as every other Anthropic customer. Truly, it&#8217;s notable that the internal structure of the Amazon and Anthropic relationship <em>didn&#8217;t</em> follow this structure to begin with, and does explain some historical oddities in how Amazonians talk about their AI use vs. basically the entire rest of the industry.</p><p>This leaves Amazon, heavily entrenched as a Claude customer, on the back foot due to three structural changes.</p><ul><li><p>Efficiency gains in serving Anthropic models will now accrue to <em>Anthropic</em>, not Amazon, as they&#8217;re moving from a structure that resembles cost-plus to instead something more akin to retail pricing.</p></li><li><p>Trainium recursion: Amazon&#8217;s chip improvements will, instead of reducing a bill that Amazon pays, improve Anthropic&#8217;s <em>margin</em> on a bill Amazon pays.</p></li><li><p>In the runup to a presumed IPO, this is effectively cleaning up a bunch of related-party exception cases. &#8220;Everyone pays us per token&#8221; is a lot cleaner structurally without an &#8220;except for this Amazon deal&#8221; outlier.</p></li></ul><p>This may well be structurally intended; we don&#8217;t know what the actual contract says or the intent of its framers. If I had to bet, I&#8217;d say that this was always the plan: in time Amazon would move to token based pricing as their usage declines. Remember, they&#8217;ve invested tens of billions into Anthropic; you don&#8217;t surprise your partners out of the blue with a completely new pricing structure unless you&#8217;re Broadcom.</p><p>The alternative is that when an Amazon spokesperson was cited as saying it was &#8220;incorrect that changes from our expanded collaboration will increase our costs,&#8221; they said it while sharpening an ostentatiously large knife. Amusing though that may be, I don&#8217;t buy it.</p><p>Curiously, a number of open weight models have already made their way into Kiro (&#8221;Amazon Basics Cursor&#8221;) in recent months. Four open model families (Qwen, Deepseek, MiniMax, and GLM; all of them behind the current frontier versions) have made their way into that dev tool&#8217;s model selector as of this writing, with Qwen3 Coder Next discounted to a shocking .05x multiplier of whatever the hell a Kiro &#8220;credit&#8221; is supposed to be. (GPT 5.6 Sol rides the other end of the pricing curve at 2.4x multiplier.)</p><h2>Pushback against the &#8220;frontier&#8221; branded token</h2><p>As per the same reporting that gave us that contract leak, Amazon has been distilling Claude internally (by all accounts their contract allows this; Amazon is not stupid, plus Anthropic has enough riding on the outcome of a legal decision here that they&#8217;re not likely to pick a ~$2.6T company as their test case), while also taking a $50B OpenAI stake as a potential second source since their &#8220;Nova&#8221; models have so far not exactly taken the world by storm. Meanwhile, &#8220;whatever Microsoft is doing&#8221; apparently includes Microsoft routing Copilot prompts to its own MAI group&#8217;s models. Microsoft&#8217;s AI CEO (which makes him sound like a model himself) Mustafa Suleyman&#8217;s stated goal of <a href="https://thenextweb.com/news/microsofts-ai-chief-says-the-company-wants-to-eliminate-what-it-pays-anthropic">zero Anthropic spend</a> makes it clear that model choice <em>for the hyperscalers themselves</em> is a critical path strategic item. All three hyperscalers have tried and failed at launching industry-leading frontier models; their path now means they have to blunt single-vendor risks, and that path increasingly looks like open models.</p><p>And the industry agrees! When you look at the plethora of great Chinese open weights models at a cost of zero beyond &#8220;the cost to actually run the models,&#8221; it&#8217;s rational to think that the model labs may start thrashing around in ways that don&#8217;t align with the goals of the hyperscalers themselves. We see the pattern starting to emerge; OpenRouter&#8217;s &#8220;top 5 by weekly tokens&#8221; leaderboard are as of this writing now all Chinese open weight models:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gTlw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gTlw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png 424w, https://substackcdn.com/image/fetch/$s_!gTlw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png 848w, https://substackcdn.com/image/fetch/$s_!gTlw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png 1272w, https://substackcdn.com/image/fetch/$s_!gTlw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gTlw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png" width="1092" height="493" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:493,&quot;width&quot;:1092,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article image" title="Article image" srcset="https://substackcdn.com/image/fetch/$s_!gTlw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png 424w, https://substackcdn.com/image/fetch/$s_!gTlw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png 848w, https://substackcdn.com/image/fetch/$s_!gTlw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png 1272w, https://substackcdn.com/image/fetch/$s_!gTlw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6373e606-8b27-4187-8fd0-5a6c029a63c0_1092x493.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Price disclosure via executive order</h2><p>A vendor meters what constrains it. (Mostly; AWS has yet to price for giving services crappy names.) Early AWS&#8217;s data flows were such that they priced for egress, but ingress cost them little once those paths were built; ergo, they had ingress bandwidth to burn so they priced it as free. What does a vendor do when it has a constraint that it&#8217;s structurally incapable of pricing? As we&#8217;re learning, they lobby about it.</p><p>The current regulatory apparatus in the US is clearly playing ball. Both Anthropic and OpenAI have been sounding the alarm about the dangers of open weight models, though the dangers seem to be largely to their own business models. The FUD machinery is in full swing, suggesting that these models will &#8220;phone home,&#8221; (trivially detectable for any reasonable infrastructure), require uploading sensitive data to the labs that created these models (these models are large files available for download; there is no need for them to be hosted by any particular provider), and might act as &#8220;sleeper agents&#8221; ready to awake and write garbage code when triggered (you hired Steven and he already does that after lunch most days).</p><p>Technical enforcement here is basically a fantasy; you can&#8217;t crack down on the downloading of these models any more than the RIAA could crack down on the downloading of MP3s decades ago, leaving their toolbox pretty empty past &#8220;criminal liability for model possession&#8221; which is a little too far down the information suppression playbook.</p><p>There&#8217;s an enforcement asymmetry here as a result. The state can raise the risk premium on Chinese weights, but not their price. On-prem inference is undetectable; you can&#8217;t really do much about it beyond policy. Ignoring first amendment issues entirely, &#8220;a government mandate to increase token spend by 20x&#8221; would go over like a lead balloon.</p><p>The entrenched are thus being taxed for depending on models that can be repriced (Amazon) or repossessed (Fable), while the models Washington is calling dangerous are showing themselves to be the only ones that can&#8217;t be either.</p><h2>What&#8217;re you to do?</h2><p>And so, barring an ability to predict the geopolitical futures, the smart enterprise instead builds in model optionality. It&#8217;s hard! I&#8217;ve found that a lot of the stuff I&#8217;ve built over the past year overindexes on Claude, and hence I&#8217;ve been sleeping on OpenAI&#8217;s excellent work&#8212;but as I&#8217;m refactoring that dependency out, I&#8217;m also planning for open models use as well. When you don&#8217;t know what rug&#8217;s about to get pulled, get several rugs and polish the linoleum underneath.</p><p>&#8212;C</p>]]></content:encoded></item><item><title><![CDATA[The Model Stopped Being the Product This Week]]></title><description><![CDATA[Five "unrelated" stories, one repricing: the value in AI has moved to everything around the model, if you're paying attention. You read this newsletter; obviously you are.]]></description><link>https://artificialconfidence.com/p/the-model-stopped-being-the-product</link><guid isPermaLink="false">https://artificialconfidence.com/p/the-model-stopped-being-the-product</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Wed, 15 Jul 2026 14:20:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LrGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F359fc899-ddd0-4b25-950b-95adb8a34ae1_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The most important thing that happened in AI this week is that nothing particularly important happened to a model. Instead, a coding agent got caught exfiltrating repositories. A lab shipped a flagship that costs more, yet holds less context than its predecessor. Apple sued OpenAI over hardware engineers. Anthropic moved a migration deadline for the third time. OpenAI repriced prompt caching. The coverage has largely treated these as five stories, but I disagree. I think they&#8217;re one story, and it&#8217;s not about any of the companies involved.</p><p>Here&#8217;s my thesis: model capability is commoditizing rapidly, and the actual competition is now around the inputs around the model. Training data, training environments, distribution endpoints, inference capacity, memory bandwidth. Each of this week&#8217;s events is the market getting ahead of that via pricing.</p><p>Remember, I spent a decade watching this exact transition happen in cloud, where the story stopped being &#8220;whose compute is better&#8221; years before the marketing did. The prices were the leading indicator there, too. Welcome to Cloud Economics, refactored for the AI era.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free because I like watching numbers go up.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2>Prices, Prices Everywhere</h2><p><strong>Development sessions got a price, and it turns out you&#8217;ve been paying it.</strong> The <a href="https://www.internationalcyberdigest.com/xais-grok-build-cli-uploads-entire-git-repositories-to-a-google-cloud-bucket/">Grok Build incident</a> is being covered as a security failure, which it absolutely is. It&#8217;s also a financial disclosure. A vendor doesn&#8217;t wire full-repository upload into a coding tool because it&#8217;s careless; that stuff is wildly expensive to move around at scale. (Do not talk to me about &#8220;free ingress,&#8221; at this scale it is oh so very far from free.) A vendor does this because authentic development sessions on real codebases are the scarce input for training coding models, and there is no synthetic substitute. I posit that this single fact neatly explains a set of economics the industry treats as separate mysteries: why coding agents are priced below cost, why vendors tolerate flat-rate plans that lose gobs of money on heavy users, why the <a href="https://x.ai/news/grok-4-5">Grok 4.5 launch post</a> presents &#8220;trained alongside Cursor&#8221; as a selling point. The &#8220;subsidy&#8221; is in fact them shrewdly making a purchase. &#8220;Your session data&#8221; is what they&#8217;re getting here, and it&#8217;s wildly valuable to them. Your coding agent is cheap for the same reason the casino comps your room: the house is not confused about who's paying. This week the price became accidentally visible, and once a price is visible, a market forms around it: expect increased disclosure, with actual-training-excluded tiers at an explicit premium, and understand that that premium is the market rate for your data. Enterprises are currently selling that asset for zero. Most of them haven&#8217;t noticed they&#8217;re the seller.</p><p>&#8220;But my vendor already says they don&#8217;t train on my data.&#8221; Sure. Read what those statements actually govern. There are three separate planes here: what gets collected off your machine, what gets retained, and what gets trained on. &#8220;We don&#8217;t train on your data&#8221; speaks to exactly one of the three, and is entirely compatible with collecting and retaining everything for &#8220;debugging,&#8221; &#8220;abuse monitoring,&#8221; or the ever-popular &#8220;service improvement.&#8221; The Grok toggle was this distinction rendered as UI: it governed training consent while the repos kept flowing. And these are attestations; you can&#8217;t verify any of them from outside the vendor&#8217;s infrastructure, which is why the &#8220;what&#8217;s on the wire&#8221; test mattered. An attestation is a pinky promise with a legal department, backed in this case by the full faith and credit of... Elon Musk. Meanwhile, notice that ZDR is something you upgrade <em>to</em>. Vendors already sell zero-retention tiers at enterprise pricing, which means the training-excluded premium I&#8217;m predicting isn&#8217;t a prediction at all, because it already exists. It&#8217;s just been filed under &#8220;compliance&#8221; instead of &#8220;the price of your data,&#8221; and nobody&#8217;s done the subtraction.</p><p><strong>Training environments became the closed layer.</strong> The same launch post contained the week&#8217;s second disclosure, because why wouldn&#8217;t it. A model trained alongside a specific tool is no longer separable from that tool; its benchmark performance is model-in-harness performance, and some fraction of the capability stays behind when you migrate harnesses. (It misses you.) This matters because it inverts the commoditization story at exactly the layer where buyers feel it. Weights leak, distill, and converge; that layer is commoditizing and will continue to do so. The RL environment and the interaction data do not leak (y&#8217;know, ideally). So the defensible asset has moved from the model to the environment the model was shaped in, and switching costs, which everyone believes are falling, are instead reconstituting themselves as capability that evaporates on migration. This is the next generation&#8217;s lock-in story. Nobody even has to sign an &#8220;all-in&#8221; agreement about it this time.</p><p><strong>Distribution endpoints became existential&#8212;in both directions.</strong> <a href="https://techcrunch.com/2026/07/10/apple-sues-openai-over-alleged-trade-secret-theft/">Apple&#8217;s suit against OpenAI</a> reads like big tech poaching drama until you notice what it&#8217;s actually about: batteries, logic boards, hardware engineers. OpenAI is building a device. Apple is defending the device layer as the crown jewels at the same moment it&#8217;s swapping Siri&#8217;s intelligence over to Google. Apple&#8217;s institutional bet is that hardware is the moat and intelligence is a commodity input you re-source annually. OpenAI&#8217;s bet is the inverse: intelligence is the moat and hardware is the escape route from platform tollbooths. Exactly one of those bets can be right, and at least one of them is currently suing like it knows it. The Gemini swap is evidence for both positions simultaneously. It proves that frontier models are now swappable components at the platform layer, and it proves that platform owners will commoditize any lab that doesn&#8217;t control an endpoint. Which yields the prediction the on-device conversation keeps missing: local inference isn&#8217;t a threat arriving at the labs, it&#8217;s a necessity arriving from inside them. You cannot ship a performant device whose intelligence round-trips to a datacenter at datacenter prices and datacenter latency. The most committed future buyers of efficient small-model inference <em>are the frontier labs themselves</em>.</p><p><strong>Inference capacity got a shiny new rationing mechanism; no it is not the deadline.</strong> Anthropic <a href="https://x.com/claudeai/status/2076351399999557669">extended included Fable 5 access</a> a third time in five weeks. The consensus read is capacity strain plus indecision, with a dash of &#8220;GPT-5.6 scared the everloving poop out of them.&#8221; That&#8230; doesn&#8217;t explain the behavior. A capacity-constrained vendor enforces a cutoff; a revenue-seeking one prices it. They&#8217;re not doing either of these. Either they&#8217;re idiots without a plan, which, sorry, I do not accept, or there&#8217;s something else going on here. Extending repeatedly, with the metered structure pre-announced and the date held loose, is a <em>measurement</em>. Each cliff-and-extension cycle produces production data on migration behavior and price elasticity before the credits pricing is committed. This isn&#8217;t about generosity or suddenly finding capacity; it&#8217;s market research on user behavior. You aren't getting a reprieve; you've been enrolled in a study. Compensation is not being offered. This is the template for how the industry exits flat-rate pricing generally: softly, instrumented to high heaven to get behavioral measurements out of it, with no pre-announcement, and the deadline drama that everyone&#8217;s focused on? Well that, gentle reader, is simply where the unsightly seam shows up. Underneath it sits the allocation logic. If inference margins are anywhere near the healthier published estimates, every subsidized subscription token carries an opportunity cost in API revenue, which makes consumer flat-rate plans permanent loss leaders held for data and mindshare purposes. The deadline keeps moving because holding to it was never the objective.</p><p><strong>Memory bandwidth got a price tag.</strong> GPT-5.6 <a href="https://openai.com/index/gpt-5-6/">reached general availability</a> with sticker prices flat across generations and one structural change: cache writes now bill at 1.25x the uncached input rate. Two things follow, and neither is &#8220;sneaky stealth price hike they&#8217;re hoping you won&#8217;t spot.&#8221; First, except for AWS&#8217;s Managed NAT Gateway data processing (itself a study in bastardry), a vendor meters what constrains it. A surcharge on cache writes is OpenAI stating, in a place that kinda sucks at spin, that memory bandwidth is its binding scarcity. As billing dimensions multiply across providers, each price sheet in turn becomes an increasingly understandable map of that provider&#8217;s infrastructure economics; if I can be slightly dismissive, the discipline of reading them is worth more than most of the analyst coverage it will eventually replace. Second, and further out: this pricing taxes statefulness precisely as long-running, memory-heavy agents become the industry&#8217;s designated growth product. The invoice-optimal agent under this structure compresses its context and recomputes at inference time rather than remembers. We have built a computer that forgets things on purpose to save money, which is the most relatable technology has been in years. Billing design is now (or will be soon) shaping agent architecture, the technically optimal system and the economically optimal system are diverging, and the discipline that bridges that gap doesn&#8217;t have a name yet. Cloud economics didn&#8217;t really have a name in 2016 either. I ended up making a career out of it anyway.</p><h2>What this predicts</h2><p>Follow those five repricings forward and they converge on a small number of claims I&#8217;m comfortable committing to.</p><p><strong>The capability conversation isn&#8217;t going to be a meaningful purchasing input.</strong> Within a couple of years, &#8220;which model is best&#8221; will sound the way &#8220;which cloud has better VMs&#8221; sounds now: a question that was settled by convergence and replaced by questions about everything surrounding the commodity. The buyers who adapt early will evaluate harness compatibility, data terms, egress behavior, and billing structure while their competitors are still reading leaderboards. You weren&#8217;t buying EC2 instances by comparing them to VMs past the first few years either.</p><p><strong>Verification displaces attestation.</strong> The Grok incident demonstrated that wire-level auditing is cheap, public, and career-making for whoever runs it. Vendor responses will bifurcate: some will sell provable boundaries, most will move execution server-side where the question &#8220;what left my machine&#8221; cannot be asked, because &#8220;all of it&#8221; is the obvious answer. Auditability itself becomes a product attribute, and the locally-executing agent, currently treated as the legacy form factor, becomes the compliance-preferred one in regulated industries just as vendors finish abandoning it.</p><p><strong>Pricing becomes the primary disclosure channel.</strong> Deadline behavior, regional availability gaps, cache surcharges, credit structures: commercial behavior is now leaking more true information about lab economics and capacity than any filing or keynote I&#8217;ve ever seen. The analytical edge shifts to whoever reads those signals systematically. Price sheets are the new S-1s, except they update monthly and nobody audits them.</p><p>And the composite prediction, the one this whole week of news drama leads me to: <strong>the labs finish becoming cloud-shaped.</strong> Consumption-driving products, vertical solutions, owned endpoints, loss-leader tiers, capacity rationing, billing complexity. Every incentive documented above pushes the same direction. You'll know it's complete when one of them throws a 60,000-person conference  about it. I've already got the name picked out: re:Inference. The invoice is in the mail. The interesting question for the next two years isn&#8217;t &#8220;which lab has the best model.&#8221; It&#8217;s which lab builds the best business around the fact that the model no longer matters most, and which vendors in the surrounding ecosystem get absorbed the way the early AWS ecosystem was.</p><p>So no, I'm not going to rank this week's news. I'm going to keep reading pricing pages, because they're the one place a vendor's marketing department doesn't get a vote. You should probably do the same.</p><p>&#8212; C</p>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence: They all picked September 1st]]></title><description><![CDATA[Fable came back last week, and Anthropic already moved its own leaving date once. Meanwhile GitHub, Google, and Anthropic all set their real price hike for the day after Labor Day, when your finance t]]></description><link>https://artificialconfidence.com/p/artificial-confidence-they-all-picked</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-they-all-picked</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Thu, 09 Jul 2026 00:02:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LrGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F359fc899-ddd0-4b25-950b-95adb8a34ae1_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve done cloud economics long enough to have a particular knee-jerk reaction: when a vendor swears nothing changed, I stop reading the announcement and start reading the pricing details. This week four of them demonstrate why this is the correct reaction.</p><p>Anthropic got its frontier model back and the first thing they did was start a countdown timer. GitHub, Google, and Anthropic all scheduled a subtly hidden price increase for the same day. Nvidia raised a sticker price 55% and didn&#8217;t schedule anything, because when you&#8217;re the only shovel store in the gold rush you don&#8217;t need a calendar. And that creepy little routing tracker somebody found in Claude Code two weeks ago vanished without ceremony.</p><p>So, strangely, nobody changed anything and yet everything changed.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">You should subscribe before your friends do for hipster cred.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><div><hr></div><h2>What actually changed (adjusted for spin)</h2><h3>Fable came home. It&#8217;ll be gone again before sunrise.</h3><p>The export controls came off June 30th and <a href="https://www.anthropic.com/news/redeploying-fable-5">Fable 5 came back July 1st</a> (presumably for reasons tied to prediction markets), ending a nineteen-day stretch where Anthropic&#8217;s best model just... wasn&#8217;t. Great. Now let&#8217;s read the fine print, because it won&#8217;t hold still long enough for me to hit the publish button. AS OF NOW, Fable is included in your plan, up to half your weekly limit, <a href="https://www.anthropic.com/news/redeploying-fable-5">originally through July 7th</a>, then dropping to usage credits at $10 and $50 per million. Then, with about a day to spare, Anthropic <a href="https://9to5mac.com/2026/07/01/claude-fable-5-cleared-to-return-as-us-lifts-anthropics-export-control-restriction/">pushed that date to July 12th</a>.</p><p>So the frontier came back on a one-week trial, and when the week was almost up, they moved the end of the week. A <a href="https://www.bleepingcomputer.com/news/artificial-intelligence/claude-fable-5-isnt-permanently-leaving-subscriptions-anthropic-says/">Claude Code engineer went on X The Everything App&#174;</a> to promise it&#8217;ll return to subscriptions &#8220;as soon as capacity allows,&#8221; which is the thing you say when everybody&#8217;s noticed the model is leaving. And that&#8217;s sort of the whole issue&#8217;s theme: this month, even Anthropic&#8217;s deadlines are on layaway.</p><h3>About that jailbreak</h3><p>The relaunch shipped a tighter safety classifier, and tighter classifiers all do the same thing: trip on your actual work. It <a href="https://www.anthropic.com/news/redeploying-fable-5">flags more ordinary coding and debugging</a>, reroutes you to Opus 4.8, and tells you it did it so you can fume. People spent the week calling the new Fable nerfed, which is what &#8220;more false positives, out of an abundance of caution&#8221; feels like when you&#8217;re the one hitting them face-first.</p><p>Now, the jailbreak that got this thing federally yanked was Amazon researchers prompting Fable into finding software vulnerabilities. Anthropic&#8217;s own testing then <a href="https://www.anthropic.com/news/redeploying-fable-5">confirmed that Opus 4.8, GPT-5.5, Kimi, and even Haiku 4.5 could find the same ones</a>. They repossessed the frontier model over a trick the cheap model does too. And Amazon, who found the bypass, is an Anthropic investor and one of the clouds Anthropic is now scrambling to switch back on. What the&#8212;just, what?</p><h3>Sonnet 5, surprise pricing next quarter</h3><p>Same morning Fable came back, Anthropic shipped <a href="https://www.anthropic.com/news/claude-sonnet-5">Sonnet 5</a>. New default on Free and Pro. Launches at $2 and $10 per million... through August 31st. After that, $3 and $15. The price they&#8217;re using to sell you the switch expires the week after Labor Day.</p><p>The pitch is &#8220;Opus quality at Sonnet money,&#8221; and to their credit <a href="https://www.anthropic.com/claude-sonnet-5-system-card">the system card admits it isn&#8217;t quite Opus</a>. Crank it to full reasoning to close the gap and, <a href="https://www.marktechpost.com/2026/06/30/anthropic-claude-sonnet-5-vs-sonnet-4-6-vs-opus-4-8-agentic-coding-benchmarks-api-pricing-and-cost-performance-tradeoffs-compared/">per day-one testing</a>, it can cost more than Opus 4.8 for the same result. The cheap model, run hard, laps the expensive one on price.</p><p>&#8220;Roughly cost-neutral&#8221; also rests on a new tokenizer that turns the same text into up to 35% more tokens. Same move they pulled at Opus 4.7, under the same &#8220;prices unchanged&#8221; banner. My favorite tell: the launch post went up with a cost chart, and within hours they <a href="https://www.anthropic.com/news/claude-sonnet-5">quietly swapped it</a> for one that assumes a ten-million-token budget per task. They fixed an &#8220;underestimate&#8221; by admitting how much it burns.</p><h3>Nvidia jacked shovel prices</h3><p>Thinking of escaping all this by buying your own hardware? Welp. Nvidia&#8217;s RTX Pro 6000 Blackwell, the 96GB card people buy specifically to run big models at home, ideally in homes with loading docks, <a href="https://www.tomshardware.com/pc-components/gpus/nvidia-raises-rtx-pro-6000-blackwell-gpu-pricing-to-usd13-250-55-percent-increase-over-msrp-in-a-years-time">now lists at $13,250</a>. A year ago it was $8,565, up 55%. Rent one in the cloud instead for around $2.43 an hour, and congratulations: you&#8217;re back on a metered clock.</p><div><hr></div><h2>Everybody circled the same Tuesday</h2><p>Here&#8217;s what you only see if you put the calendars side by side, which is kinda the entire job of this newsletter:</p><p>Sonnet 5&#8217;s intro price ends and it jumps to $3/$15 on September 1st.</p><p>GitHub Copilot went to token metering June 1st and kept every sticker price flat. The promo credits that hide the sting for Business and Enterprise <a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/">run June through August</a>, then evaporate, once again, on September 1st. (I covered the switch itself in June. The new part is the first bills landing and the credits running out.)</p><p>Google got bored again and renamed Vertex AI to &#8220;Gemini Enterprise Agent Platform,&#8221; and in the fine print of the new pricing page, a stack of fresh agent meters switch on by a now-recurring date. That&#8217;s right, <a href="https://cloud.google.com/vertex-ai/pricing">Memory Bank and Sessions billing both start September 1st</a>. The same rename also &#8220;pulled an AWS redefining &#8216;Serverless&#8217;&#8221; and killed scale-to-zero, so your idle nodes now bill around the clock. That&#8217;s a price increase they don&#8217;t have to meaningfully disclose up front, which feels just terrific as a customer.</p><p>Nobody raised a <em>headline</em> number. Even so, they all set the real pricing time bombs to go off the week you&#8217;re at the beach and your finance team is out. GitHub even framed the whole thing as <a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/">&#8220;plan prices aren&#8217;t changing,&#8221;</a> which is true the same way that &#8220;I didn&#8217;t eat the last cookie, I relocated it to my mouth&#8221; is true.</p><div><hr></div><h2>The tracker that silently disappeared</h2><p>Somebody found that Claude Code, when you point it at a non-Anthropic endpoint, was sneakily judging where your requests were headed and hiding the results of that judgement in the apostrophe of the &#8220;Today&#8217;s date&#8221; line. Anthropic, the Ethical AI Company&#174; keyed it to a hidden, lightly-scrambled list of Chinese cloud domains and Claude resellers. I <a href="https://registry.npmjs.org/@anthropic-ai/claude-code/-/claude-code-2.1.91.tgz">pulled the package and reproduced it</a>: 147 domains, four different apostrophes, lightly obscured, a timezone check for Shanghai and Urumqi. Such transparent. Very safe. Much enterprise trust.</p><p>I checked again this week, because &#8220;it&#8217;s still there&#8221; changes with time, just like Anthropic changes previously announced dates. It&#8217;s not there anymore. The date line is back to a plain apostrophe, and the scrambled list doesn&#8217;t decode anywhere in the current build. Sometime in the last few days it just... left. No changelog, no &#8220;we heard you,&#8221; no mea culpa unless I missed a freaking tweet, nothing. That&#8217;s rapidly becoming the Anthropic house style. When they throttled Claude Code&#8217;s reasoning in the spring, they later called it <a href="https://www.anthropic.com/engineering/april-23-postmortem">&#8220;the wrong tradeoff.&#8221;</a> When they nerfed Fable at launch, they told Wired they <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">&#8220;made the wrong trade-off.&#8221;</a> That apology phrasing has become a reusable component now and can be bound to a keyboard shortcut. The hidden tracker tracker didn&#8217;t even rate that milquetoast treatment, it just stopped being there.</p><p>I won&#8217;t call it spyware, because it technically wasn&#8217;t. It never directly phoned home. Point Claude Code at Bedrock and the little marker rode into Amazon&#8217;s datacenter and (if you can trust AWS, which in this sense I do) it died there, useless. It only really did anything when a reseller relayed your prompt back to Anthropic itself. Catching resellers is fair. Doing it by sewing a classifier into a piece of punctuation and scrambling the list so it won&#8217;t show up in a search is not. You don&#8217;t hide the part you&#8217;re proud of.</p><p>And look at the deal Anthropic just signed to get Fable back: <a href="https://www.anthropic.com/news/redeploying-fable-5">proactively detect security risks, coordinate with the government on releases, report malicious activity</a>. It describes almost exactly what their tool was already doing in secret. I&#8217;d have more sympathy for the distillation panic driving all this if the product ringing the alarm weren&#8217;t &#8220;the entire internet distilled,&#8221; with a class action&#8217;s worth of other people&#8217;s books stirred in for fuller body and better mouth feel. You don&#8217;t get to install a secret tripwire against scraping and clutch your pearls when scraping is your entire company.</p><div><hr></div><h2>One last thing</h2><p>You have to read the fine print when budgeting, because everyone&#8217;s hiding material changes these days. Fable&#8217;s yours until July 12th, presumably. Sonnet gets pricier September 1st. Copilot&#8217;s discount dies September 1st. Google&#8217;s new meters wake up September 1st. And the one reasonably <em>honest</em> thing anybody shipped all month, a tracker hidden in a curly quote, got memory-holed.</p><p>Put September 1st in your calendar before it shows up on your invoice, because I promise you it will. Then go count your own tokens, because apparently nobody selling them to you is going to do it for you.</p><p>See you next week.</p><p>&#8212; C</p>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence: Ask permission, not forgiveness]]></title><description><![CDATA[Anthropic shipped Fable 5 to the public; we know what happened next. OpenAI cleared GPT-5.6 with Washington before launch and got to ship. This bodes ill.]]></description><link>https://artificialconfidence.com/p/artificial-confidence-ask-permission</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-ask-permission</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Tue, 30 Jun 2026 14:31:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LrGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F359fc899-ddd0-4b25-950b-95adb8a34ae1_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;m at the AIE World&#8217;s Fair this week in San Francisco. Hit reply and say hi if you&#8217;re around; I&#8217;m looking to meet some of you folks.</p><div><hr></div><p>Anthropic put Fable 5 in front of the public and then promptly got pantsed by the White House. OpenAI <a href="https://openai.com/index/previewing-gpt-5-6-sol/">previewed GPT-5.6</a> last week as a release it would stage &#8220;in phases per the federal government&#8217;s request,&#8221; having broken out the requisite crayons, puppets, and juice boxes to walk Washington through the model.</p><p>It would appear that the question of &#8220;which frontier lab is winning&#8221; will be settled less by technology than by whom can speak more effectively to political power.</p><p>And for downtime enthusiasts, <a href="https://azure.microsoft.com/en-us/blog/claude-in-microsoft-foundry-is-now-generally-available/">Microsoft Foundry now hosts Claude models</a>.</p><div><hr></div><h2>What Actually Changed (Adjusted For Spin)</h2><h3>Mythos 5 came back for a select few, Fable 5 still MIA</h3><p>On June 26, Commerce Secretary Howard Lutnick cleared Anthropic to redeploy Mythos 5 to a defined cohort: somewhere around a hundred critical-infrastructure organizations on an approved annex, their foreign-national employees, and a slice of civilian government. Everyone else still needs an export license, which... is a sentence that should probably change how you read a vendor contract. Fable 5, the general-purpose model that we mere mortals were using, stays dark going on three weeks, with Pentagon and NSA sign-off still &#8220;pending&#8221; like it&#8217;s a CloudFormation update.</p><p>The price of the partial thaw, per the reporting on Lutnick&#8217;s letter, was Anthropic committing to work with the government on &#8220;protocols and standards and releases,&#8221; plus a clause letting the Secretary revoke the access list whenever he likes. That is the regulatory equivalent of getting your car out of impound on the condition that you subscribe to the impound lot&#8217;s newsletter.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe to <em>my</em> newsletter, I&#8217;ll never put your car into impound.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p>For the record, I pay Anthropic every month and I am locked out of Fable along with the rest of the civilian population. The fallback is Opus 4.8, which is fine the way a rental car is fine, right up until you reach for a button that isn&#8217;t there. I did seven hours of road tripping last week in a rental car; forgive me if my analogies are as crappy as that Chrysler minivan.</p><h3>GPT-5.6 shipped, with the government holding the leash</h3><p>OpenAI previewed GPT-5.6 the same afternoon: three tiers (Sol at $5/$30 per million tokens, Terra at $2.50/$15, Luna at $1/$6), real model IDs, a system card rating it high-capability in cybersecurity, and a new &#8220;ultra&#8221; mode that throws subagents (a fancy approximation for &#8220;money&#8221;) at hard problems. The catch is, to no great surprise, the access terms. Available via API and Codex only, limited to &#8220;trusted partners,&#8221; with the government signing off on customers case by case during the preview. OpenAI&#8217;s own post says it doesn&#8217;t believe this should become the long-term default, a notably principled stance from the company that&#8217;s currently demonstrating how well the default works when you&#8217;re the one who pre-cleared. So... points to them for ideological consistency.</p><div><hr></div><h2>The gift of gab becomes the gift of lab</h2><p>Anthropic&#8217;s sin, in the eyes of the people who can turn its models off, was launching to the public and only discovering where the line was afterward. OpenAI&#8217;s move was to find the line first, in private, and ship inside it. This reads as a memo to every frontier lab: going to market first is now the <em>liability,</em> and pre-clearing with the US government is this cycle&#8217;s moat.</p><p>Which is why both companies have abruptly discovered enthusiasm for codifying the process, with an August deadline looming on the executive order that set up voluntary model vetting. They want different things here. Anthropic wants rules because their absence cost it three weeks (so far!) and its flagship. OpenAI wants rules because it just proved it&#8217;s good at the part that happens before the rules, the relationship part, and would like that game written down while it&#8217;s ahead. Whoever&#8217;s best at the clearance dance gets to choreograph it, and right now that&#8217;s not the lab with the higher benchmark.</p><p>For those of us who who are cloud economists and not the armchair equivalent, this bolts a column onto the model-selection spreadsheet that I&#8217;ve never seen on any pricing page. &#8220;Cost-per-correct-answer&#8221; means precisely nothing the instant a Cabinet secretary can dark the model, and the vendor most likely to still be serving requests next quarter is apparently the one best at managing Washington, not the one with the tightest unit economics. The open-weight tier I went on about last week (GLM 5.2 and Kimi and Cohere&#8217;s small one) doesn&#8217;t touch frontier cyber capability&#8212;and isn&#8217;t pretending to. But a download has no approved-entity list to get cut from, and &#8220;merely very good and impossible to repossess&#8221; is aging into the responsible adult&#8217;s hedge.</p><p>The future is stupid.</p><div><hr></div><h2>One last thing</h2><p>Add that horrible column. Before you commit a workload to any hosted frontier model, price the tokens, then ask who has standing to switch it off and how good your vendor is at staying on that person&#8217;s good side. The bill was apparently the <em>easy</em> part. Availability is now&#8230; political. Some would say it always has been.</p><p>See you next week.</p><p>&#8212; C</p>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence: You can’t repossess a download]]></title><description><![CDATA[A government nastygram took the country&#8217;s best coding model offline for everyone. What arrived the same week isn&#8217;t a frontier replacement; it&#8217;s an Opus-class workhorse you can own.]]></description><link>https://artificialconfidence.com/p/artificial-confidence-you-cant-repossess</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-you-cant-repossess</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Tue, 23 Jun 2026 15:12:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LrGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F359fc899-ddd0-4b25-950b-95adb8a34ae1_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;m in Indianapolis this week to keynote the <a href="https://www.midwestcommunityday.com/">AWS Community Day</a> tomorrow morning; come by if you&#8217;re in town!</p><div><hr></div><p>Last week&#8217;s thesis was that everything in AI is rented, and rented things get repossessed. This happens via a pricing email, a status page, or apparently a Commerce Department directive that lands at 5:21pm on a Friday and pulls Fable 5 and Mythos 5 offline for everyone because someone failed to genuflect properly to their feudal lord. Eleven days later they&#8217;re still dark, and the company says that&#8217;ll get fixed &#8220;in the coming days.&#8221; I am not going to rehash this.</p><p>This week&#8217;s about what the market did while the frontier sat unplugged and Twitter sat grousing.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe so actually thoughtful takes start living in your inbox.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><p></p><div><hr></div><h2>What Actually Changed (Adjusted For Spin)</h2><p>Three open-weight coding models landed since last I (graced|darkened) your inbox. Zhipu&#8217;s (ghesundheit) <a href="https://Z.ai">Z.ai</a> <a href="https://z.ai/blog/glm-5.2">shipped GLM 5.2 on June 13</a>, live on its Coding Plan day one, offering a million-token context window, with MIT weights to follow. Moonshot <a href="https://huggingface.co/moonshotai/Kimi-K2.7-Code">shipped Kimi K2.7-Code on June 12</a>, a trillion-parameter model under a modified MIT license at $0.95 per million input tokens. Cohere&#8217;s <a href="https://cohere.com/blog/north-mini-code">North Mini Code</a> arrived June 9, Apache 2.0, 30 billion parameters. That last one matters specifically because Cohere is Canadian, which cuts against the ongoing &#8220;China bad&#8221; narrative. Zhipu, meanwhile, has been on the <a href="https://www.federalregister.gov/documents/2025/01/16/2025-00704/addition-of-entities-to-and-revision-of-entry-on-the-entity-list">Commerce Entity List since January 2025</a>: right now the industry cares a hell of a lot more that this model hasn&#8217;t been ripped away from them. Now we <em>know</em> that models can be turned off at the will of the US government. For some use cases that&#8217;s unacceptable, and as we know by now the internet treats censorship as breakage and routes around it; this is going to lead us to fascinating places.</p><p>We should be clear about what this trio of models represents. None of them are Fable-class; that tier is exactly what got pulled, and nothing open touches it&#8212;yet! But I ran GLM 5.2 this week, served through <a href="https://baseten.co">Baseten</a>, against the multi-file work I&#8217;d normally hand Opus 4.8, and it&#8217;s an Opus-class contender: it did the job, it didn&#8217;t need four tries, and nobody had to approve my access once I shoved the key into my harness. Fable, for the brief window I had access, was pretty clearly going to be the big, slow, excellent frontier model you reached for occasionally. Opus has positioned itself as the model you reach for all day, and the all-day tier now has an open-weight peer that lives on hardware no angry White House letter can reach.</p><div><hr></div><h2>Follow The Money (It Went To The Landlord)</h2><p>The capital people noticed before you did. Baseten, the platform I ran that model through, <a href="https://techcrunch.com/2026/06/18/ai-inference-startup-baseten-reportedly-raising-1-5b-months-after-its-last-mega-round/">closed a $1.5 billion round yesterday</a> at a valuation of up to $13 billion. &#8220;Up to,&#8221; because it&#8217;s split-priced at $11 billion for some investors and $13 billion for others, because of reasons that aren&#8217;t worth going into. That roughly triples its $5 billion mark from January, on something like $600 million of annualized revenue. Baseten doesn&#8217;t make a model. It makes the unglamorous layer that turns a free download into something that answers in production, across clouds it doesn&#8217;t own. Right now that&#8217;s looking like the most fundable pitch in AI: not the model (whose financials are, let&#8217;s say... dubious), but rather the place you run the model once you&#8217;ve decided it should be something nobody upstream can switch off.</p><p>For the record, the company that just had two models repossessed is the one I pay every month, so weigh my enthusiasm for ownable weights accordingly.</p><div><hr></div><h2>One last thing</h2><p>A rented model can be turned off by someone who isn&#8217;t your vendor; we all answer to the sovereign government that controls the places we sit. But weights on your own disk are a lot more durable, and can change jurisdiction very quickly. So the number worth chasing is starting to look less like which hosted frontier tops the leaderboard, but rather what a correct answer costs on the version you host yourself, where &#8220;they turned it off&#8221; isn&#8217;t a meaningful risk factor. We&#8217;re building toward a real figure on that. Price your exit before you need it.</p><p>See you next week.</p><p>&#8212; C</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free because I am a goddamned delight.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[No failover for the Commerce Secretary]]></title><description><![CDATA[The US government switched off Anthropic's two best models for every customer on Friday evening. Your disaster-recovery plan has a runbook for an outage, yet nothing for a letter from Commerce.]]></description><link>https://artificialconfidence.com/p/no-failover-for-the-commerce-secretary</link><guid isPermaLink="false">https://artificialconfidence.com/p/no-failover-for-the-commerce-secretary</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Mon, 15 Jun 2026 14:47:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xoFp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xoFp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xoFp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png 424w, https://substackcdn.com/image/fetch/$s_!xoFp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png 848w, https://substackcdn.com/image/fetch/$s_!xoFp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!xoFp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xoFp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Focused image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Focused image" title="Focused image" srcset="https://substackcdn.com/image/fetch/$s_!xoFp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png 424w, https://substackcdn.com/image/fetch/$s_!xoFp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png 848w, https://substackcdn.com/image/fetch/$s_!xoFp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!xoFp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15a43e1c-2457-49e9-8f39-91b384d27dbb_2048x1152.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The problem with writing this newsletter is that very often the things that are true when I start writing are no longer true an hour or so later, when I&#8217;ve finished writing. Let&#8217;s see if I can get this out early, in-flight to the AWS Summit in NY, before the situation changes. If events once again outpace me, it&#8217;s imperative that you not email me about it. Let&#8217;s race the clock:</p><p>I&#8217;ve spent many years telling folks that various things (this week&#8217;s: AI dependency) distills down to a vendor-risk problem, with knowable failure modes: prices go up, rate limits get chokingly tight, the version or model you&#8217;ve standardized on gets Googled, the green of the status page becomes a comforting lie, etc. There are playbooks for all of these, and you can pay your way free of most of them.</p><p>On Friday evening we saw a fifth failure mode show up and the playbooks were either lacking or non-existent. The US government, or the closest thing we&#8217;ve got in this era, told Anthropic to take its two best models away from foreign nationals, Anthropic (rightly) determined it couldn&#8217;t do that selectively without risking jail time for its execs, and so it <a href="https://www.anthropic.com/news/fable-mythos-access">took them away from everyone</a>. If you were building on Fable 5 at 5:20pm you were not building on Fable at 6:05pm, much to your surprise. There wasn&#8217;t a migration window, no announced timeline, just the equivalent &#8220;how quickly can we rip the power cord out the back of these GPU rigs like we&#8217;re ripstarting some very expensive lawnmowers.&#8221;</p><p>So this week, the theme is &#8220;learning the thing you rent can not just be repriced, but can also in fact be repossessed.&#8221;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free to remind me I don&#8217;t need to charge money to sponsors.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2>What Actually Changed (Adjusted For Spin)</h2><h3>Anthropic&#8217;s best models went dark by federal letter</h3><p>Anthropic <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">released Fable 5 and Mythos 5</a> on June 9, with Fable pitched as the first time the company put a model of that tier in front of the general public. The Commerce Department, in a directive signed by Secretary Lutnick, <a href="https://www.anthropic.com/news/fable-mythos-access">had some thoughts</a> the next evening at 5:21pm ET, of the form &#8220;now you listen here.&#8221; It cited national security authorities and ordered access suspended for any foreign national, whether inside or outside the country, including Anthropic&#8217;s own foreign-national employees. Because there is no clean and prison-proof way to wall those people off one model at a time, Anthropic disabled both models for every customer to stay compliant. It&#8217;s been reported that this is the first time a leading lab has pulled a deployed model offline because the federal government said so, and I have not found a counterexample. Given the reaction we&#8217;re seeing, I kinda think I would have heard if it were otherwise.</p><p>The stated reason for this is a jailbreak. As per Anthropic, the demonstrated technique amounts to asking the model to read a codebase and fix the flaws in it, which surfaced a few minor vulnerabilities that other public models find too. That seems less a &#8220;jailbreak&#8221; than it is &#8220;the exact capability they&#8217;re using to market the model.&#8221; In effect, it seems that &#8220;being good at the job&#8221; is now a reason your vendor can lose the ability to offer it to you.</p><p>Anthropic is complying, because prison, while saying that it believes this to be a misunderstanding, CNN, citing Axios, reports the directive would require a license not just for export but for <a href="https://www.cnn.com/2026/06/13/business/anthropic-mythos-model-national-security">domestic transfer of the models</a>, which, if this remains true, means moving the weights between two datacenters in the same country is now a paperwork event. Yay, paperwork!</p><p>Either way, the practitioner takeaway is not &#8220;the government overreached&#8221; or &#8220;Anthropic was reckless.&#8221; Pick whichever of those you like at dinner, fine, whatever, I really don&#8217;t care that much about the reason. Because the takeaway is that your disaster-recovery plan has a runbook for us-east-1 falling over due to coup or asteroid strike or whatever, and nothing whatsoever for a one-page letter from Commerce, and no amount of multi-region anything fixes that, because the thing that failed was not the infrastructure unless you believe in the 8th layer of the OSI model (<em>politics</em>).</p><h1>Meanwhile, three coding tools metered the buffet in two weeks</h1><p>While the dramatic repossession was getting the headlines and Twitter screaming, a far quieter one has been going on relatively stealthily, and it is the one that will actually move your bill. The flat all-you-can-eat seat for agentic coding is being retired across the category, because an agent that runs autonomously for an hour burns serious compute and the flat fee was subsidizing the heavy users out of the light ones&#8217; pockets, as is always the way with subscriptions.</p><p>We&#8217;ve talked about GitHub Copilot doing this a couple of weeks ago, but I missed that <a href="https://devin.ai/blog/windsurf-is-now-devin-desktop/">Windsurf became Devin Desktop on June 2</a> in an over-the-air update from Cognition, with Cascade, the local agent a lot of CI pipelines call by name, going end-of-life July 1. This wasn&#8217;t coordinated, so put your tinfoil hats away; it&#8217;s an emergence of what we&#8217;re seeing industry wide. What it means is &#8220;if you have automation that invokes Cascade, that is a serious deadline with your name splattered on it, and &#8220;the editor renamed itself overnight and my pipeline broke&#8221; is a sentence you would presumably prefer not to say to your team on July 2.</p><p>The honest read is that metered pricing is more correct than flat pricing. A heavy user and a light user genuinely do not cost the vendor the same, and pretending otherwise was always a subsidy with an expiration date. The catch is the one every cloud customer already knows in their soul: &#8220;you only pay for what you use&#8221; is a wonderful promise right up until an agent decides to use rather a lot of it at three in the morning.</p><h2>Reliability: A Brief Retrospective</h2><h2>Follow The Money (Or Watch It Follow Itself)</h2><p>SpaceX <a href="https://www.npr.org/2026/06/12/nx-s1-5855004/stock-ai-spacex-ipo-elon-musk">went public on June 12 under SPCX</a>, raised about $75 billion in the largest IPO ever recorded, priced at $135, and <a href="https://www.cnbc.com/2026/06/12/spacex-ipo-spcx-live-updates.html">closed its first day up 19% near $161</a>, which put the whole thing somewhere north of two trillion dollars, brushing against Amazon&#8217;s own valuation intraday. This is of course an AI story now, because the future is stupid, but also because in February SpaceX absorbed xAI and folded it into an AI division, and that division reported an operating loss of $6.36 billion last year and burned another $2.47 billion in the first quarter of this one. The market looked at a rocket company carrying an AI furnace and apparently decided it was worth more than all but a handful of companies that have ever existed. This seems fine,</p><p>Note: only about 4% of the shares are actually trading; insiders are locked up for six months and Musk holds something like 85% of the voting power. So the two-trillion-dollar number you are reading is the price the world is paying for the 4% it is allowed to touch, multiplied across the 96% it is not, which is both a fine way to generate an enormous headline and a poor way to learn what the company is worth. Morningstar&#8217;s discounted-cash-flow model lands around $780 billion, and the difference between that and two trillion is the dollar value of believing the burn turns into something. (I am pointedly not mentioning the formal investigations now open in several jurisdictions over Grok, because nothing about the substance of those is a joke and I am not going to treat it as one.)</p><div><hr></div><h2>One last thing</h2><p>Go look at your architecture diagram and find the box, or be honest, boxes where the model lives. You have probably already drawn the arrows for what happens when it times out, and maybe even what happens when the price changes. Add the case where it is simply gone on a Friday because someone in Washington sent a letter, decide now whether your business survives that week, and price the answer accordingly. The vendors spent two years teaching us to pay only for what we use. This week we learned &#8220;&#8230;only for as long as we are allowed to use it.&#8221;</p><p>See you next week, unless events once again outpace me..</p><p>&#8212; C</p>]]></content:encoded></item><item><title><![CDATA[Claude Opus / Fable / Shitpost]]></title><description><![CDATA[In April it was too dangerous to release. In June it's in your Pro subscription until the 22nd, then it isn't, then maybe it is again.]]></description><link>https://artificialconfidence.com/p/claude-opus-fable-shitpost</link><guid isPermaLink="false">https://artificialconfidence.com/p/claude-opus-fable-shitpost</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Tue, 09 Jun 2026 21:45:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!66pc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!66pc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!66pc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png 424w, https://substackcdn.com/image/fetch/$s_!66pc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png 848w, https://substackcdn.com/image/fetch/$s_!66pc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png 1272w, https://substackcdn.com/image/fetch/$s_!66pc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!66pc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png" width="1456" height="977" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:977,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Full size image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Full size image" title="Full size image" srcset="https://substackcdn.com/image/fetch/$s_!66pc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png 424w, https://substackcdn.com/image/fetch/$s_!66pc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png 848w, https://substackcdn.com/image/fetch/$s_!66pc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png 1272w, https://substackcdn.com/image/fetch/$s_!66pc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F946a3fc7-1aa3-4d15-853d-adff4b3d32ca_2528x1696.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>It figures; I whack &#8220;Send&#8221; and a new model drops. Let&#8217;s see here...</p><p>In April, Anthropic told the world that Mythos was too dangerous to release. It was so scary! But it&#8217;s great. But it&#8217;s scary! The fear-based marketing spiel was... a bit much. They built a government consortium around keeping it in a vault, handed it to a hundred and fifty cyberdefenders under the codename Project Glasswing, and said they had no plans to ship it to the public. The model was so good at finding software vulnerabilities that letting strangers use it got framed as a national-security question.</p><p>And then two months later they shipped it to anyone with a Pro subscription.</p><p>The thing they shipped is called Fable 5, following their naming convention of &#8220;types of writing.&#8221; One day I look forward to seeing them release Claude Shitpost, but that day is not today. Fable is the <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">same underlying model as Mythos</a> with a layer of classifiers bolted on top that watch what you ask and, roughly one session in twenty, decide you can&#8217;t have the good model and route your question to the previous one instead. So the most capable model Anthropic has ever made generally available also ships with an asterisk that occasionally hands you the model it replaces. Let&#8217;s talk about the asterisk, because it&#8217;s where all the interesting economics live.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free before my token bill wrecks my month.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2>The price is Opus, in a hurry</h2><p>Fable 5 is <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">out today on the API and every paid tier</a> at $10 per million input tokens and $50 per million output. That&#8217;s exactly double Opus 4.8&#8217;s $5 and $25. It is also, to the dollar, the price of Opus 4.8 in Fast Mode. So the new frontier model costs precisely what the old frontier model costs when you ask it to type faster.</p><p>The capability claims are the usual launch buffet, and some of them are even real. Stripe says it <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">compressed months of engineering into days</a> on a fifty-million-line Ruby codebase, which is the kind of specific, customer-attributed claim worth more than ten benchmark charts. It also beat Pok&#233;mon FireRed using only screenshots, which is delightful and tells you approximately nothing about your bill.</p><h2>The classifiers route you to Opus 4.8, and they tell you</h2><p>Here&#8217;s the part to understand before you wire it into anything. Fable ships with classifiers covering three buckets: cybersecurity, biology and chemistry, and distillation, which is their word for people trying to clone the model&#8217;s capabilities to train a competitor. Trip one and the <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">response comes from Opus 4.8 instead</a>, with a note telling you it happened.</p><p>Anthropic says this fires in <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">under 5% of sessions</a> and that a fallback beats a refusal, which is true. But analyze the economics for a second. You&#8217;re paying Fable&#8217;s price, $10 and $50. When the classifier fires you get an Opus 4.8 answer, the same Opus you could have bought directly for $5 and $25. So one session in twenty, the safety system&#8217;s net effect is to charge you double for the cheaper model, presumably while you watch the little note explain why. The biology and chemistry net is currently cast wide enough that <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">most requests</a> in those areas fall back, so if your work touches a beaker you&#8217;re paying frontier rates for the floor below frontier a good deal more than one time in twenty.</p><p>This isn&#8217;t me objecting to the safeguards. The uplift case for a model that finds zero-days across every major OS is a concern and they&#8217;re right to be nervous about it. Rather, it&#8217;s an observation about who&#8217;s holding the meter when the safeguard does its job, because it feels somehow wrong to refuse to do something but take someone&#8217;s money for it anyway.</p><h2>Thirty-day retention, now mandatory</h2><p>Anthropic is also requiring <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">30-day data retention on all Mythos-class traffic</a>, business customers included, across first- and third-party surfaces. They won&#8217;t train on it, they&#8217;re logging human access to it, they delete it after the window in almost all cases, and we can probably pretend this isn&#8217;t happening, right? The stated purpose is catching multi-request attacks and reducing false positives, which is plausible.</p><p>It&#8217;s also a thing that, until today, your enterprise data agreement may have promised you didn&#8217;t have to do. If you&#8217;re the person who negotiated zero-retention into a contract so you could put a frontier model in front of regulated workloads, &#8220;mandatory 30-day retention on the new model tier&#8221; is a sentence you want to read before, rather than after, the migration. Enterprise legal departments will almost certainly scream about this before inevitably signing the contract anyway.</p><h2>The subscription that comes and goes</h2><p>Now the part that reads like a hostage note with somehow worse legibility.</p><p>Fable 5 is <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">included on Pro, Max, Team, and seat-based Enterprise plans through June 22</a>. On June 23 it comes off those plans, and using it requires usage credits. After that, &#8220;when sufficient capacity allows,&#8221; they aim to put it back into the subscriptions. They&#8217;ll extend the included window if capacity allows, and restore standard access as quickly as they can.</p><p>Read in order, the offer is: it&#8217;s free, then it costs money, then it&#8217;s free again, terms and dates subject to how the GPUs are feeling. This is a thirteen-day trial that they&#8217;re calling a launch, attached to a model they&#8217;re describing as the most capable they&#8217;ve ever released. The honest version of the sentence is &#8220;we don&#8217;t have enough compute to give everyone this model, so we&#8217;re going to give it to everyone for two weeks and then take it back,&#8221; and to their credit the announcement nearly says exactly that, in the register of a company that would prefer you focus on the giving rather than the taking-back. It pointedly does not address how, if capacity is so dear, they&#8217;re able to give it to everyone for those two weeks at launch, when interest and thus demand is clearly spiking.</p><p>I read the rollout schedule three times to make sure I had the sequence right. I believe that I did. It&#8217;s in, then out, then in, and the only firm date in the whole arrangement is the one where it leaves.</p><p>It&#8217;s that &#8220;demand is hard to predict&#8221; and &#8220;capacity allows&#8221; are included in a launch announcement for a flagship model, which is the AI-industry version of a restaurant putting its best dish on the menu with a footnote reading &#8220;if we feel like it.&#8221;</p><p>So if you&#8217;re putting Fable into production, do the boring thing before June 23: figure out which of your workloads trip the classifiers, because those are the ones where you&#8217;re paying double for Opus 4.8, and price your fallback path at the model you actually get rather than the one on the label. The capability is, as per early reports, real. The asterisks are painfully real. Plan accordingly.</p>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence: xAI, the Neocloud]]></title><description><![CDATA[xAI raised frontier-lab money, then used it to quietly turn into a GPU landlord, renting its Memphis data center to the two competitors out-shipping it. Apple conceded it can&#8217;t host its own AI either.]]></description><link>https://artificialconfidence.com/p/artificial-confidence-xai-the-neocloud</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-xai-the-neocloud</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Tue, 09 Jun 2026 16:18:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K5CL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K5CL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K5CL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png 424w, https://substackcdn.com/image/fetch/$s_!K5CL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png 848w, https://substackcdn.com/image/fetch/$s_!K5CL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png 1272w, https://substackcdn.com/image/fetch/$s_!K5CL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K5CL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png" width="1456" height="977" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:977,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated image" title="Generated image" srcset="https://substackcdn.com/image/fetch/$s_!K5CL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png 424w, https://substackcdn.com/image/fetch/$s_!K5CL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png 848w, https://substackcdn.com/image/fetch/$s_!K5CL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png 1272w, https://substackcdn.com/image/fetch/$s_!K5CL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81549d1a-7cbb-440d-8199-019e51e6eb6e_2528x1696.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There&#8217;s an indicator that shows up in an S-1 when a company&#8217;s stopped doing the thing it raised money to do, and SpaceX&#8217;s prospectus has it. xAI raised the kind of capital you raise to win the AI model race and be the Bestest Boy, but then spent the last week disclosing that its actual business is instead renting GPUs to the companies who&#8217;re winning the model race. Meanwhile, Apple conceded it cannot run its own assistant on its own servers and is paying Google a billion dollars a year to borrow one. And Anthropic, which rents xAI&#8217;s spare compute to keep Claude running, filed to go public on revenue it won&#8217;t show you yet, then watched Claude fall over twice in four days.</p><p>So somehow we&#8217;ve arrived here, a place where everyone in AI is now renting the one part of the stack they spent three years insisting they had to own themselves.</p><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe to feed my ever-growing need for external validation.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2>What Actually Changed (Adjusted For Spin)</h2><h3>Google agreed to pay SpaceX $920 million a month for compute</h3><p>In a <a href="https://techcrunch.com/2026/06/05/google-will-pay-spacex-920m-per-month-for-compute/">June 5 filing</a>, Google committed to roughly $920 million (what&#8217;s a few million here or there between friends?) a month through June 2029 to rent capacity from SpaceX. The world&#8217;s largest owner of AI compute cannot build data centers fast enough to keep up with itself, so it is renting them from a rocket company. Totally normal. Very sane. This is all fine.</p><h3>Apple shipped the betas that let you replace Siri</h3><p>iOS 27, iPadOS 27, and macOS 27 developer betas landed yesterday with the new Extensions framework: you can now designate Claude, ChatGPT, or Gemini as the default provider for Apple Intelligence features. Apple also deprecated SiriKit in favor of App Intents, the only framework that will talk to the rebuilt assistant because Apple is Very Special. GA in the fall. A real API and availability change, not a keynote promise&#8212;but can we really trust Apple&#8217;s keynotes after their Apple Intelligence oversteps?</p><h3>Claude went down. Twice.</h3><p>June 2 and June 5. Oops. #hugops to them.</p><div><hr></div><p></p><h2>xAI Is A Neocloud Now (They Just Can&#8217;t Say So)</h2><p>xAI built Colossus 1, the Memphis data center fueled by gas turbines and bad faith, to train Grok. Then, per SpaceX&#8217;s S-1, it decided the smarter move was to <a href="https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/">rent the whole thing to Anthropic</a> for $1.25 billion a month through 2029, and last week added Google at another $920 million. Combined, that&#8217;s about $26 billion a year from compute it bought to do something else. Anthropic raised Claude&#8217;s usage limits the day that deal was announced, the clearest sign the constraint was always the compute, never the demand.</p><p>That&#8217;s the part the prospectus dances around with the same care you&#8217;d bring to defusing a landmine. SpaceX calls this a &#8220;dual monetization strategy&#8221; and says it &#8220;allows us to monetize unused compute capacity,&#8221; which is the corporate-finance way of describing the spare bedroom you rent out as a &#8220;diversified hospitality vertical.&#8221; It&#8217;s unused because Grok didn&#8217;t need it, presumably because most businesses have minimal use cases in the workplace for &#8220;revenge porn.&#8221; xAI moved its real training to Colossus 2. So model lab is now a landlord, the tenants are the companies beating it, and this timeline remains profoundly stupid.</p><p>The counterfactual everyone politely skips is &#8220;why not point all that compute at making Grok better.&#8221; The answer&#8217;s in the public record: the last time Grok made headlines for its capabilities, it was enthusiastically introducing itself as MechaHitler. Thirty billion dollars of GPUs is not, right now, the obvious value-maximizing play, so the GPUs are for rent and the S-1 found a much nicer word for it.</p><p>This matters because of what trades Friday: SPCX opens June 12 at a fixed $135, around a $1.75 trillion valuation, roughly 95 times last year&#8217;s revenue and the largest IPO ever attempted. A chunk of what&#8217;s being sold as frontier-AI upside is, upon inspection, a leasing business. Because nobody can build anything in this industry without sleeping with everybody else, Google (newly signed as a Colossus tenant) was an early SpaceX investor, holds a stake, and has a director on the board. So one of the two customers anchoring xAI&#8217;s revenue narrative also profits when that narrative prices the IPO higher. What an economist calls vertical integration, and what everyone else calls a massive conflict of interest.</p><div><hr></div><p></p><h2>Everyone Else Is Renting Too</h2><p>If xAI is the supply side of the great AI sublet, Apple spent yesterday as the demand side. Tim Cook&#8217;s last keynote as CEO unveiled a rebuilt Siri that runs on a custom 1.2-trillion-parameter Google Gemini model, for which Apple is reportedly paying around a billion dollars a year (which is, to Apple, chump change). The company that designs its own silicon and built a retail religion on owning the whole stack apparently could not, as it turns out, build the one part that now matters to everyone.</p><p>It gets better. Apple <a href="https://www.macrumors.com/2026/06/04/apple-siri-rely-on-google-nvidia-chips/">originally tried to host the model on Private Cloud Compute</a> and found, per The Information, that a trillion-parameter model ran too slowly at Siri&#8217;s scale. So the heaviest queries route to Nvidia B200s, and the contract leans on Nvidia&#8217;s on-chip encryption to keep Google from reading them. The &#8220;we own the whole stack&#8221; company is now shipping as their flagship announcement what is in effect a group project, which feels surreal.</p><p>So Claude and ChatGPT lost the &#8220;which AI lab does Apple partner with&#8221; bake-off, but the consolation prize is arguably better. Extensions let you set them as the default for the rest of Apple Intelligence.</p><p>And of course, none of this is free of history. Apple is shipping (in beta) the contextual Siri it promised at WWDC 2024 and didn&#8217;t deliver, a gap that cost it a $250 million settlement whose approval hearing lands ~nine days from now. If they got it right this time, then the features will finally arrive, as someone else&#8217;s model on rented hardware, two years and a class action later. Which, in the current AI industry, resembles a... passing grade?</p><p></p><div><hr></div><h2>Reliability: A Brief Retrospective</h2><p>Anthropic filed its confidential S-1 on June 1. The next day, Claude went down. Three days later, it went down again, taking claude.ai, the API, Claude Code, and Cowork with it. Oops.</p><p>I have a professional interest in Claude staying online (because I am *NOT* going to write IAM policies myself like some kind of agrarian farmer), and watching the tool you use to do your job blink out twice in one week, while its parent is valued like &#8220;every power utility combined,&#8221; concentrates the mind on the gap between calling yourself a utility and behaving like one. You don&#8217;t wonder if water is going to come out of the tap when you turn it on, unless you&#8217;ve been vibe-plumbing again.</p><p>The customers who ran Claude through Vertex or Bedrock mostly rode it out, which is the lesson nobody selling you a trillion-dollar single point of failure wants underlined. It&#8217;s priced as critical infrastructure and yet it&#8217;s run, for now, like a startup having a week. If you&#8217;re putting it in production, architect for the Tuesday it isn&#8217;t there.</p><div><hr></div><p></p><h2>One last thing</h2><p>Every deal this week is a bet that the tokens keep flowing through somebody else&#8217;s building. xAI rents out the data center it couldn&#8217;t train on, Apple rents the model it couldn&#8217;t build, Google rents capacity from a rocket company, and Anthropic rents all of the above, files to go public on the strength of it, and then trips and falls down the availability stairs. A trillion dollars of valuation are currently resting on the premise that inference is something you have to drive somewhere to buy.</p><p>A Stanford lab spent <a href="https://arxiv.org/abs/2511.07885">last November</a> documenting that the laptop already in your bag handles something like nine of every ten everyday questions on its own, and got five times better at it in two years. Nobody&#8217;s neocloud is priced for the day the easy tokens stop making the trip.</p><p>So watch the intelligence-per-watt curve, not the IPO calendar. The tokens are already walking home.</p><p>See you next week.</p><p>&#8212; C</p>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence: GitHub Repriced the Habit it Built]]></title><description><![CDATA[GitHub evolved its billing model, and you&#8217;ll feel it soon. Everyone else annualized their best month and called it revenue.]]></description><link>https://artificialconfidence.com/p/artificial-confidence-github-repriced</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-github-repriced</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Fri, 05 Jun 2026 00:44:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gpUB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHJViewebAAE1uVB.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I was overcome by events at Microsoft Build + fwd:CloudSec this week, so I&#8217;m sending this later than I normally do. I missed you, too. Now then:</p><p>GitHub moved Copilot to usage-based billing on June 1, and the folks who responded in the first day are precisely the heavy agentic users GitHub built this pricing structure to find. Meanwhile, back at the ranch, everyone else spent the week reporting run-rates: Cognition annualized a number, Anthropic presumably has one inside a confidential filing, and a small chorus of VCs sang out what I had assumed was already widely known: ARR means &#8220;a strong month, multiplied by twelve, assuming number only ever go up.&#8221; The net result of this is that the one dollar figure that actually changed got buried under a pile of imaginary ones instead.</p><div><hr></div><h2>What Actually Changed (Adjusted For Spin)</h2><h3>GitHub changes course</h3><p><a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/">GitHub Copilot moved to usage-based billing on June 1</a>. Premium request units are gone, replaced by token-metered &#8220;AI Credits,&#8221; which while sounding inscrutable, isn&#8217;t THAT far removed from the ever-shifting definition of a token. The base subscription prices did not change, which GitHub would very much like you to notice given that they haven&#8217;t stopped harping on that particular detail. What did change is that those prices now describe how much you get <em>before</em> the meter starts, which... may not be how many customers would like this story to end.</p><p>Two details buried in the notes are significant here. <a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/">The fallback model is gone</a>, so when your credits run out you no longer downgrade to a cheaper model, you simply stop. It is the first subscription I have encountered that ends mid-sentence, much as I am tempted to do to this one. And a Copilot code review now <a href="https://github.com/orgs/community/discussions/192948">bills against AI Credits and GitHub Actions minutes at the same time</a>, which feels... unfortunate.</p><p>The screenshots going around show projected bills jumping from &#8220;$50&#8221; to &#8220;several thousand,&#8221; and the caveat is that those are extrapolations from people a single day into a billing period that has not yet produced a real invoice, with zero changes to their workflows. The funnier and also truer point is who is doing the extrapolating. This may be hard for some folks to hear, but you absolutely do not arrive at a terrifying projected number by accident; you encounter them after the fact, in your bill, when you&#8217;re contemplating doing something truly desperate but also cannot afford rope. You get there by being precisely the high-volume agentic user GitHub spent two years <strong>encouraging you to become</strong>, and then doing the math. The loudest complaints this week are a confession of exactly the usage profile the new pricing was built to locate. GitHub will correctly tell you your bill is atypical compared to a hypothetical spherical cow / customer. They&#8217;re learning that &#8220;correct&#8221; is not the same as &#8220;reassuring,&#8221; as it&#8217;s becoming clear that while customers value transparency, it&#8217;s as a means to the end of what they really want: predictability.</p><h3>Cursor doubled a price and let everyone watch GitHub instead</h3><p><a href="https://cursor.com/blog/composer-2-5">Composer 2.5&#8217;s Fast tier went from $1.50/$7.50 to $3.00/$15.00 per million tokens</a>, a 100% increase that landed two weeks ago and that almost nobody filed as a price hike, because it was positioned as &#8220;the more important number to go up is the version number of the model.&#8221; It&#8217;s interesting, because this is directionally the same move that GitHub made, at roughly the same time, to half the outrage&#8212;because a new model number is a press release, while a new price is relegated to a footnote.</p><p>And the introductory rates are expiring in a chorus. Composer 2.5&#8217;s launch promo ended May 25, Codex Pro&#8217;s ended May 31, and the Opus 4.7 multiplier inside Copilot already doubled on April 30. The pattern is now established: ship at a subsidized rate, train the workflow, then let the meter find its level. If you have not locked in your workflow&#8217;s economics before the promo expires, surprise! You have some thinking to do.</p><div><hr></div><h2>Follow The Money (Or Watch It Follow Itself)</h2><h3>Cognition raised $1B, and you should read the metrics carefully</h3><p><a href="https://techcrunch.com/2026/05/27/ai-coding-startup-cognition-raises-1b-at-25b-pre-money-valuation/">Cognition closed a Series D of more than $1 billion at a $26 billion post-money valuation on May 27</a>, up from $10.2 billion eight months earlier. The headline statistic, repeated everywhere, is that 89% of the code committed at Cognition is now written by Devin, the company&#8217;s own AI software engineer and is totally not just some guy in a trench coat and a fake moustache.</p><p>What this means, filtered through my snarky lens, is that a company that sells an AI software engineer is reporting that its AI software engineer writes most of its software, and offering this as proof the product works. It is the cleanest available example to date of a vendor grading its own homework, on a test it wrote, in a classroom it owns, then issuing a press release about the score. The 89% may be entirely real, but it&#8217;s also the least independent benchmark imaginable until next week, when something will no doubt surpass it somehow.</p><p>The number that underpins that is revenue that grew from $37 million to $492 million in twelve months, a roughly 53x multiple on the valuation. To Cognition&#8217;s credit, they pulled an Andy Jassy and called it run-rate, which is the honest term. The problem comes in the shape of everyone who read &#8220;run-rate&#8221; and heard &#8220;revenue.&#8221;</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/cognition/status/2059660758531940856&quot;,&quot;full_text&quot;:&quot;1/ We&#8217;ve raised over $1B at a $26B valuation, led by <span class=\&quot;tweet-fake-link\&quot;>@Lux_Capital</span>, <span class=\&quot;tweet-fake-link\&quot;>@generalcatalyst</span>, and <span class=\&quot;tweet-fake-link\&quot;>@8vc</span>.\n\nOur enterprise usage has grown &amp;gt;10x since the start of this year, and our run-rate revenue grew to $492 M.\n\nWe launched Devin two years ago as the first AI software engineer. Since &quot;,&quot;username&quot;:&quot;cognition&quot;,&quot;name&quot;:&quot;Cognition&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1765909640364068865/MvH-m0gd_normal.jpg&quot;,&quot;date&quot;:&quot;2026-05-27T15:39:26.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HJViewebAAE1uVB.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/k99LLLyWhZ&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:165,&quot;retweet_count&quot;:200,&quot;like_count&quot;:2467,&quot;impression_count&quot;:856528,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><h3></h3><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe so I feel like somebody&#8217;s actually reading this</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h3>Everyone&#8217;s revenue is a run-rate now, and a few VCs finally said so</h3><p>ARR used to mean annual recurring revenue: money a customer was contractually obligated to pay you. Think &#8220;I have signed up for a one year contract with you at a fixed fee schedule, committing me to pay you $X a month.&#8221; <a href="https://techcrunch.com/2026/05/22/how-vcs-and-founders-use-inflated-arr-to-kingmake-ai-startups/">In a TechCrunch piece last month</a>, Spellbook CEO Scott Stevenson called the current usage of the term a &#8220;scam,&#8221; and he is not even slightly wrong about the mechanics. Usage-based billing breaks the &#8220;contracted&#8221; half of ARR, so a strong month gets annualized&#8212;much like a salesperson who will take their best ever month, multiply it by 12, and claim that was their total annual compensation.</p><p>The really damning line came from an unnamed investor in the same piece: once one company in a category does it, the rest nearly have to, just to keep pace. That is a prisoner&#8217;s dilemma of revenue reporting brought to you by someone funding both prisoners.</p><h3>Anthropic filed the most confident empty document of the year</h3><p><a href="https://techcrunch.com/2026/06/01/anthropic-files-to-go-public/">Anthropic confidentially filed a draft S-1 on June 1</a>, beating OpenAI to the announcement that carries the least checkable information of any in the AI IPO cycle. OpenAI will likely shortly file an &#8220;S-1o&#8221; with improved reasoning or whatnot. &#8220;Confidential&#8221; means, obviously, that we cannot read a word of it. The valuation and the roughly $47 billion run-rate everyone is quoting come from the last funding round and what the company told its investors, not from the audited document that is currently sealed from view. This is the filing that has real numbers of the &#8220;make these up and you may well serve prison time&#8221; variety.</p><div><hr></div><h2>The Agents Got Expensive. They Did Not Get Safer.</h2><p>So let&#8217;s review the week. The bill for agentic coding went up (GitHub, Cursor). The valuation of agentic coding went up (Cognition, $26 billion). And the independently measured quality of what these agents actually commit did not move because of course it didn&#8217;t.</p><p><a href="https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report">CodeRabbit&#8217;s analysis of 470 real-world pull requests found AI-co-authored code introduced up to 2.74x more security vulnerabilities</a> than human-written code. <a href="https://www.veracode.com/blog/genai-code-security-report/">Veracode tested more than 100 models and found 45% of AI-generated samples introduced an OWASP Top 10 vulnerability</a>, a pass rate that has not improved across testing cycles despite a steady stream of vendor claims that the latest model finally fixed it. The security curve (motto: &#8220;the one nobody puts on a slide&#8221;) is remaining flat, regardless of how the capability advances.</p><p>Which puts Cognition&#8217;s 89% in a stark light: if your AI writes 89% of your code, and independent measurement says AI-written code ships vulnerabilities at multiples of the human rate, then 89% is less a productivity statistic than it is a description of your attack surface&#8212;annualized.</p><p>Meanwhile, inference itself keeps getting cheaper. <a href="https://api-docs.deepseek.com/quick_start/pricing">DeepSeek&#8217;s V4-Flash runs around $0.14 per million input tokens</a> against GPT-5.5&#8217;s $5.00, a gap of roughly 36x on input and north of 100x on output, at comparable performance on many tasks. So the raw cost of intelligence is collapsing in public while the cost of the tools you actually code in went up this week. So the &#8220;token cost is shrinking&#8221; is offset by &#8220;so let&#8217;s immediately burn as many as they can in nondeterministic ways in the harnesses.&#8221;</p><div><hr></div><h2>One last thing</h2><p>If you run AI workloads, do the boring thing this week: pull your own usage numbers before the next promo expiry (there&#8217;s always another one!), and find out which of your workflows is hanging out in a cohort waiting to be repriced. The vendors already know. Remember: the only number in this entire issue you can fully verify is the one on your own invoice.</p><p>See you next week.</p><p>&#8212; C</p>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence: Banned by Pentagon, blessed by Pope, paid for by you]]></title><description><![CDATA[Five assumptions your AI vendor stack lost in five business days. The trade press will not be on the cc line of your Q3 invoice.]]></description><link>https://artificialconfidence.com/p/artificial-confidence-banned-by-pentagon</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-banned-by-pentagon</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Tue, 26 May 2026 16:58:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LrGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F359fc899-ddd0-4b25-950b-95adb8a34ae1_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1>Artificial Confidence: Banned by Pentagon, blessed by Pope, paid for by you</h1><p>I have spent a fairly depressing decade reading AWS bills for a living, and the dependable feature of every quarter remains constant: the vendor&#8217;s headline announcements and the customer&#8217;s resulting bill invariably have nothing to do with each other. The AI vendors are improving on that formula in realtime, and it&#8217;s really something to see.</p><p>This week alone: three trillion-dollar IPOs queued in the same fortnight, a $30 billion Series H closing, two cybersecurity products launched in the same news cycle, an SDK supplier acquihired and wound down, a 235-page papal encyclical personally presented by a pope for the first time in modern history, and a new payments protocol with a transaction class called, on purpose, &#8220;Human Not Present.&#8221;</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free, because the AI industry is hiring more PR firms than the unsubscribe button can keep up with.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The press is reading this as a string of capability and capital announcements. Good for them; here in the real world the customers paying these vendors&#8217; bills just lost roughly five assumptions they didn&#8217;t know they had. Most of them are only gonna find out from their Q3 invoice, but you&#8217;re ahead; count all five below.</p><div><hr></div><h2>What Actually Changed (Adjusted For Spin)</h2><h3>Anthropic bought the company that generated SDKs for OpenAI and Google</h3><p>Anthropic <a href="https://www.anthropic.com/news/anthropic-acquires-stainless">acquired Stainless on May 18</a> for <a href="https://www.uctoday.com/productivity-automation/anthropic-acquires-stainless-the-sdk-and-mcp-startup-behind-every-official-claude-api-library/">a reported $300 million-plus</a>, roughly double Stainless&#8217;s December valuation. Stainless powered every official Anthropic SDK, and also generated client libraries for OpenAI, Google, and Cloudflare. Anthropic <a href="https://techcrunch.com/2026/05/18/anthropic-has-acquired-the-dev-tools-startup-used-by-openai-google-and-cloudflare/">is winding down all hosted Stainless products</a> which is how we know it&#8217;s an acquihire. <strong>You assumed the official Claude or OpenAI SDK you </strong><code>pip install</code><strong>ed was a first-party product</strong>, by which I mean that you didn&#8217;t consider where the SDK came from at all, because why would you? And also the vendor never told you. Surprise, that assumption was incorrect. Isn&#8217;t it great when we get to learn things together? Instead, it was a third-party deliverable from a vendor Anthropic now owns and is sunsetting. If your stack relied on Stainless&#8217;s hosted workflow, you have an engineering project this quarter that you did not have last week, and the vendor management group at OpenAI has the same one. Maybe you can trade tips?</p><h3>Project Glasswing&#8217;s first month found 10,000 critical vulnerabilities</h3><p>Anthropic <a href="https://interestingengineering.com/ai-robotics/anthropic-project-glasswing-10000-software-vulnerabilities">published the first numerical results</a> on May 22 from a fascinating project: roughly fifty partners ran Claude Mythos Preview against critical software for a month. The headline was over 10,000 high-or-critical-severity vulnerabilities, including <a href="https://hothardware.com/news/anthropic-project-glasswing-targets-safer-ai-agents">2,000 at Cloudflare</a> (400 high or critical), 271 patched in Firefox 150 (ten times what a comparable Opus 4 scan found in Firefox 148), and <a href="https://cybersecuritynews.com/anthropics-claude-mythos-preview-0-days/">a 90.6% true-positive rate</a> on the open-source subset that external firms verified. <strong>You assumed your security budget was priced against the scarcity of capable vulnerability scanners.</strong> That scarcity ended on May 22. Tenable, CrowdStrike, and Palo Alto Networks did not price their products against Opus 4.7 inference. They will need to, in the next RFP cycle.</p><h3>Claude Security launched the same afternoon, helpfully</h3><p>Anthropic <a href="https://hothardware.com/news/anthropic-project-glasswing-targets-safer-ai-agents">launched Claude Security in public beta</a> on May 22, an Opus 4.7-powered codebase scanner that proposes patches. It had already patched 2,100 enterprise vulnerabilities in the preceding three weeks. The company that announced an industry-wide discovery-rate problem in the morning was selling the cleanup contract by the afternoon, which is a sequencing move that should be appreciated for its craft, if not its decorum. Because that&#8217;s more than a bit crass.</p><h3>OpenAI confidentially filed S-1 the same week its Q1 margins leaked</h3><p>OpenAI <a href="https://www.cnbc.com/2026/05/20/openai-ipo-filing.html">confidentially filed S-1</a> on Friday May 22 with Goldman Sachs and Morgan Stanley, targeting a Q4 listing at $852B&#8211;$1T. The same news cycle, <a href="https://www.wheresyoured.at/news-openai-had-a-negative-122-operating-margin-in-q1-2026-and-chatgpt-growth-has-stalled/">The Information</a> reported that OpenAI generated $5.7 billion in Q1 revenue at a non-GAAP adjusted operating margin of negative 122%, meaning the company lost $1.22 for every dollar of revenue it brought in. ChatGPT weekly actives stalled at 905 million, down from a 920 million February peak; the free-to-paid conversion rate is approximately 6%. What&#8217;s that mean for you? Simply that <strong>you assumed your OpenAI rate card was sustainable.</strong> It is apparently instead subsidized by venture capital at approximately 45% of cost. The S-1 disclosure cycle is the mechanism that&#8217;s fated to end the subsidy. This is a data point for the growing thesis that the relatively cheap tokens you are using today are not going to be cheap tokens in eighteen months.</p><h3>Anthropic&#8217;s Series H closing this week at $900B-plus</h3><p>Bloomberg <a href="https://www.bloomberg.com/news/articles/2026-05-22/anthropic-to-close-over-30-billion-round-as-soon-as-next-week">confirmed</a> Anthropic&#8217;s $30 billion Series H closing next week at a pre-money valuation above $900 billion. This is the company&#8217;s second $30 billion round of 2026; the February Series G closed at $380 billion post-money. Anthropic&#8217;s reported annualized run rate moved from $14 billion in February to $45 billion in early May. Even Anthropic appears to be a little startled by Anthropic. <strong>You assumed Anthropic&#8217;s pricing discipline was a structural property of the company</strong>, the &#8220;we&#8217;re not OpenAI&#8221; pitch made flesh, as it were. At $900 billion-plus closing this week and <a href="https://www.techtimes.com/articles/317066/20260523/anthropic-funding-round-top-30b-900b-valuation-would-surpass-openai-most-valuable-ai-startup.htm">a Q4 2026 listing reportedly targeted</a>, the same public-market disclosure cycle that is about to retire OpenAI&#8217;s rate-card subsidy is now starting at Anthropic. The &#8220;disciplined alternative&#8221; pitch has an IPO clock on it, and the clock is running.</p><h3>&#8220;Human Not Present&#8221; payments became a real schema in a real standards body</h3><p>Google <a href="https://blog.google/products-and-platforms/platforms/google-pay/agent-payments-protocol-fido-alliance/">donated AP2 to the FIDO Alliance</a> and shipped v0.2, introducing autonomous-transaction support that the protocol&#8217;s own documentation <a href="https://ap2-protocol.org/ap2/specification/">officially calls &#8220;Human Not Present&#8221; payments</a>. Sixty partner organizations including PayPal, Mastercard, Visa, and American Express. The payments industry has used &#8220;Card Not Present&#8221; for two decades as the elevated-fraud category for online card use. &#8220;Human Not Present&#8221; is the same naming convention, except now the variable that has been removed from the loop is the buyer. <strong>You assumed &#8220;Human Present&#8221; was the only transaction class your e-commerce stack had to support.</strong> That assumption retired on the same Tuesday Google announced the Universal Cart.</p><div><hr></div><h2>The Pope Showed Up. The Pentagon Was Conspicuously Absent.</h2><p>This is the week&#8217;s most colorful item and consequently its least practical, which is why I am putting it after the line items rather than ahead of them.</p><p>On Monday morning in Rome, Pope Leo XIV personally presented <em><a href="https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.html">Magnifica Humanitas</a></em>, his 235-page first encyclical, the first to be personally presented by a pope in modern history. He invited one of Anthropic&#8217;s approximately forty co-founders (Chris Olah) to sit next to him and speak as one of five named presenters alongside three cardinals (the religious figures, not the birds; I&#8217;m not Simon Willison) and two theologians. The encyclical text criticizes &#8220;concentration of power and data in the hands of so few people in the private sector&#8221; and does not name a company. The Vatican did not need to. Olah was <a href="https://www.bnnbloomberg.ca/business/technology/2026/05/25/anthropics-olah-says-ai-must-be-guided-from-outside-big-tech/">confirmed by Reuters</a> as the only Big Tech representative invited to the event. Cardinal Czerny, still not a bird, asked whether Anthropic&#8217;s reputation as a safety-forward AI company had influenced the Vatican&#8217;s decision, <a href="https://www.cbsnews.com/news/pope-leo-ai-encyclical-artificial-intelligence/">said</a> &#8220;I&#8217;m sure it did,&#8221; and then added, &#8220;We dialogue with anyone. We don&#8217;t endorse.&#8221; Endorsement is not the technical term I would reach for either, but the photograph runs on every Catholic news service in five languages.</p><p>On March 3, <a href="https://www.washingtonpost.com/technology/2026/03/09/anthropic-lawsuit-pentagon/">the Pentagon designated Anthropic a national security supply chain risk</a>, <a href="https://www.cnbc.com/2026/04/08/anthropic-pentagon-court-ruling-supply-chain-risk.html">the first American company</a> ever to receive a label historically reserved for foreign adversaries, after Anthropic refused to remove guardrails on autonomous-weapons and domestic-surveillance use of Claude. The president ordered federal agencies off Anthropic via Truth Social, his social network that answers the question &#8220;what if Twitter were somehow worse.&#8221; The legal fight is, of course, ongoing.</p><p><strong>You assumed your AI vendor selection was a technical and economic decision.</strong> Picking your primary AI vendor is no longer like picking a cloud provider. It is more like picking a defense contractor: the political conditions under which you are permitted to use them are now part of the contract, and the political conditions change with the news cycle. Multi-vendor strategy is no longer a redundancy hedge, but also a political-volatility hedge. The engineering cost of building a vendor-abstraction layer that lets you swap from Anthropic to OpenAI on twelve hours&#8217; notice just became non-optional, and you can put it in next year&#8217;s budget under &#8220;geopolitical risk.&#8221; The trouble is, the models and tooling around them are differentiating, so that twelve hour window is growing longer by the day. Talk to your engineering teams about that.</p><div><hr></div><h2>One last thing</h2><p>The vendors will continue announcing capabilities and the press will continue covering those capabilities, but the important part for you is that regardless of those, the bills will continue arriving. The vendors and the press will not be on the cc line when you read your bill. Read your rate card <strong>this</strong> quarter. Understand what it&#8217;s telling you, which is harder than it looks. Audit your dependency graph. Assume the political climate around your primary vendor is different next quarter than it is this one.</p><p>See you next week.</p><p>&#8212; C</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free. I read the S-1s, the status pages, and the pricing diffs so you don't have to, and then I make fun of them on a schedule.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence: The Spec for the Agent-Native Cloud, and Who Might Actually Ship It]]></title><description><![CDATA[I tweeted a twelve-point spec for what an agent-native cloud actually needs to look like. Vercel volunteered. Cloudflare's engineers got to work. Here's the test.]]></description><link>https://artificialconfidence.com/p/artificial-confidence-the-spec-for</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-the-spec-for</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Wed, 20 May 2026 18:09:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LrGY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F359fc899-ddd0-4b25-950b-95adb8a34ae1_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I&#8217;ve spent a decade watching companies blundering into the discovery that the cloud they built their company on is not the cloud they need anymore, usually at a moment when the bill arrives or the auditor shows up. We&#8217;re heading into one of those moments now, and the existing hyperscalers are not going to be the ones who build the cloud that agents actually want to run on.</p><p>Last week, Vercel CEO Guillermo Rauch was <a href="https://x.com/rauchg/status/2055491454307582454">posting about Grok CLI deploying to Vercel</a>, and I <a href="https://x.com/QuinnyPig/status/2055492278169510299">replied</a> that an agent-native cloud platform was coming. It might be Cloudflare, it might be Vercel, and it absolutely wasn&#8217;t going to be AWS. Rauch responded inside an hour: &#8220;It&#8217;ll be &#9650;. Would love your feedback. This is our primary focus!&#8221; That seemed like a sufficiently confident claim to warrant taking him up on it, so I posted a twelve-point thread laying out the spec a serious contender would have to clear:</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free. I read the S-1s, the status pages, and the pricing diffs so you don't have to, and then I make fun of them on a schedule.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/QuinnyPig/status/2055497559813304735&quot;,&quot;full_text&quot;:&quot;Been thinking about what an \&quot;agent-native cloud\&quot; actually needs to look like. Mentioned this, and <span class=\&quot;tweet-fake-link\&quot;>@vercel</span>'s CEO replied that it'll be them. Cool! Here's the spec they (or <span class=\&quot;tweet-fake-link\&quot;>@Cloudflare</span>, or some startup not yet invented) actually have to hit. \n\nIt won't be <span class=\&quot;tweet-fake-link\&quot;>@awscloud</span>.\n\nThread...&quot;,&quot;username&quot;:&quot;QuinnyPig&quot;,&quot;name&quot;:&quot;Corey Quinn&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1840839119037218817/3aPpjjwH_normal.jpg&quot;,&quot;date&quot;:&quot;2026-05-16T03:56:22.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;@QuinnyPig It'll be &#9650;. Would love your feedback. This is our primary focus!&quot;,&quot;username&quot;:&quot;rauchg&quot;,&quot;name&quot;:&quot;Guillermo Rauch&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1783856060249595904/8TfcCN0r_normal.jpg&quot;},&quot;reply_count&quot;:28,&quot;retweet_count&quot;:36,&quot;like_count&quot;:438,&quot;impression_count&quot;:132661,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:false}" data-component-name="Twitter2ToDOM"></div><p></p><p>Cloudflare&#8217;s response came from a different altitude: <a href="https://x.com/chatsidhartha/status/2055827076897226887">Principal Systems Engineer Sid Chatterjee replied</a> &#8220;I saw your list on the thread. We&#8217;re on it. Will report back once they&#8217;re all in.&#8221; Rauch volunteered the company as the test subject; Chatterjee volunteered to do the work and produce the receipts. Both responses are legitimate, but the engineer-level &#8220;report back once they&#8217;re all in&#8221; is the one that the rest of this post is calling for.</p><p>Here&#8217;s the spec, expanded from the thread, with credit to the practitioners who supplied the parts I missed. It assumes a specific deployment shape: an agent running semi-autonomously, taking actions over minutes or hours, against real infrastructure, with real money attached. If you&#8217;re typing prompts and watching every step, you don&#8217;t need an agent-native cloud; you need a less hostile CLI. The hard but compelling work is what changes when the agent runs unattended and the platform has to be trustworthy enough that you don&#8217;t have to babysit the bill.</p><p><a href="https://x.com/rtheoryxyz/status/2056012016611905753">Robertus on Twitter</a> put the thesis better than I did: &#8220;agent-native cloud needs boring primitives more than magic. identity, permissions, logs, rollback, and cost controls before the sci-fi layer.&#8221; The vendors who lose this race build the sci-fi layer first and the boring primitives never. The vendors who win recognize that &#8220;boring primitives&#8221; is a euphemism for &#8220;the hard infrastructure problems that took AWS twenty years to get most of the way through, which is why a clean-slate competitor has a real opening.&#8221;</p><h2>Identity and blast radius</h2><p><strong>Agents need their own identity.</strong> Today every agent action is laundered through the human&#8217;s IAM role. The audit log reads &#8220;corey@duckbill did this&#8221; when the truth is &#8220;Claude&#8217;s third retry at 2am did this.&#8221; That isn&#8217;t an audit log so much as it is compliance theater. First-class agent identities have to be scoped, time-limited, and revocable, so that when a postmortem rolls around, the answer to &#8220;which agent, what session, what tools, what action&#8221; is in the log rather than reconstructed from inference and the meeting notes of whoever was on call that night.</p><p><strong>Blast radius as a primitive.</strong> &#8220;This session may spend up to X dollars, touch up to N resources, in environment Y, expiring in 30 minutes.&#8221; Today every agent is either fully privileged or fully fenced off, and the entire interesting design space is in between. Almost nobody is building there, because it requires answering hard questions about resource-graph traversal that AWS has spent a decade pretending IAM was solving.</p><p><strong>Secrets brokering.</strong> Stop making the agent fish for API keys every time it wants to light up a new service. The platform holds the secret; the agent gets a handle; calls go through the broker. A compromised agent cannot exfiltrate what it never had. This is a solved problem in OAuth flows for human users and a completely unsolved problem for agent-to-agent service calls, mostly because nobody has wanted to solve it.</p><h2>The money problem</h2><p><strong>Hard budget caps that actually halt.</strong> Not the AWS approach of &#8220;we noticed you spent $47,000 yesterday, here&#8217;s a CloudWatch email,&#8221; which is a postmortem with the dollar amount filled in, not a functioning budget control. Fail closed at the boundary. A Lambda stuck in a loop racking up data transfer charges is a real failure mode and deserves real boundary enforcement, not retroactive grief. The platform that ships caps that actually halt eliminates a category of incident I&#8217;ve spent a decade collecting war stories about: agent runs a recursive S3 list against a misconfigured bucket for fourteen hours, discovery happens at invoice time, blame routes to the most junior engineer who touched IAM that quarter.</p><p><strong>Cost circuit breakers with human escalation.</strong> The agent session has an allotment; when it depletes faster than expected, the platform pages a human to authorize more or kill it. Finding out at the end of the month is how the surprise-bill incidents keep happening.</p><p><strong>Cost preview as a first-class API.</strong> Before any state-changing call: &#8220;this adds approximately $340 per month fixed, plus $0.09 per thousand requests.&#8221; Most pricing is usage-based now, so the preview has to model the workload rather than return a single number. Agents are bad at AWS pricing because AWS pricing is bad at being prices. The platform that ships a working cost preview API breaks a fifteen-year stalemate in cloud finops.</p><h2>The reversibility problem</h2><p><strong>Gated changes by default.</strong> The agent does not mutate production directly. It opens a PR, kicks off an Action, proposes a change that a human or another agent reviews. The pattern is established and agents haven&#8217;t started routing around it; the platform&#8217;s job is to make it the path of least resistance.</p><p><strong>Time travel by default.</strong> Every state change is reversible for some defined window. &#8220;Roll back the last twenty minutes&#8221; is one command, not an archaeological dig through CloudTrail that ends with restoring yesterday&#8217;s snapshot and losing four hours of customer data in the process of un-doing the agent&#8217;s mistake. This may well be impossible to engineer for anything beyond trivial levels of complexity, but by god, we need a better answer than today&#8217;s &#8220;hope your backups are ready for an impromptu test!&#8221;</p><p><strong>Error messages designed for an LLM to act on.</strong> Not &#8220;AccessDenied: User arn:aws:... not authorized because no identity-based policy allows...&#8221; which is technically information but functionally a puzzle the agent will fail at solving. More like: &#8220;denied: this agent lacks dynamodb:Query on the &#8216;users&#8217; table; the owner can grant it at [link].&#8221; Errors as instructions, not riddles. The industry-wide bill for inference cost burned decoding AWS error messages is already in the eight figures, distributed across millions of individual agent loops where nobody is going to notice it until someone like me writes a report explaining what they&#8217;ve been paying for.</p><h2>The interface problem</h2><p><strong>The API has to be consistent.</strong> AWS has 347 services, pending an update by AWS Corporate Comms (good job, buddy! You&#8217;re making a difference here!), depending on what counts as a service this week and whether you&#8217;re counting the ones that have been deprecated but not removed from the console. Roughly 43 of them do approximately the same thing, with bespoke verbs, inconsistent pagination, regional quirks, and conventions that exist because a single Principal Engineer in 2014 had strong opinions that accidentally became load-bearing. Agents inherit this inconsistency tax at a higher rate than humans do, paying it on every retrieval against a token budget that gets spent trying to remember whether this particular service uses NextToken, pageToken, or Marker.</p><p><strong>Observability that ties action to reasoning to cost.</strong> Not &#8220;Lambda X fired&#8221; but &#8220;agent invoked Lambda X while attempting task Y, prompted by request Z, costing $0.0003 against a $5 session budget.&#8221; The AI-native equivalent of <code>dmesg</code> for distributed systems. The vendor that ships this becomes the default observability layer for agentic infrastructure, which is to say, becomes Datadog with a four-year head start. Ideally with a more dignified mascot situation.</p><p><strong>Convention over configuration as an iron rule.</strong> AWS forces explicit decisions on a thousand things with one obviously-right answer 95% of the time. The agent-native platform should have opinionated defaults; when it does need to ask, ask the human, not flail through alone burning tokens on guesses. Vercel understands this; Framework-defined Infrastructure is exactly that thesis applied to web apps. Whether it generalizes from &#8220;deploy a Next.js app&#8221; to &#8220;operate a stateful multi-agent system&#8221; is the question on which the bet rests.</p><h2>The thirteenth and fourteenth items (which I missed)</h2><p><a href="https://x.com/Hey_ross/status/2055502833043325035">Ross Brown replied</a> with what should have been item 13: a universal context injection system that pre-loads agents with the relevant architecture, the active alert posture, and the company policy (&#8221;agents may not modify the billing table without dual approval&#8221;) rather than relying on stuffing it into a <a href="https://CLAUDE.md">CLAUDE.md</a> and praying. <a href="https://usewire.io/">Wire</a> is one of the companies building this; there will be others.</p><p>Item 14, which several practitioners flagged: every spec item above is about the <em>build phase</em>, where the agent operates against the platform. The <em>run phase</em>, where what the agent built has to serve real users, is a separate set of problems the platform should solve so the agent doesn&#8217;t have to invent them, badly, from a half-remembered StackOverflow post about JWT handling. Authentication, session management, password reset flows, OAuth, MFA. The cleanest vendor example shipping on this axis is <a href="https://exe.dev/">exe.dev</a>, which puts an IAM proxy in front of every VM by default: TLS, DNS, and auth handled at the platform layer, not retrofitted into the agent-generated app. A full run-phase spec is its own post, but for now: nobody should be allowed to claim &#8220;agent-native cloud&#8221; while only solving the build-phase problems, even though the build-phase problems are the ones currently getting all the attention.</p><h2>Who&#8217;s credibly in the race</h2><p>Within forty-eight hours of the thread, my mentions filled up with founders explaining that they had already built three of these, were working on the next four, were 90% there, were the obvious frontrunner, were the only serious contender, and had also been doing this for years before anyone else noticed. The actual contenders, sorted by capital behind the claim rather than enthusiasm: Vercel and Cloudflare at the top, with meaningfully different architectural bets (Vercel: serverless functions with durable workflows; Cloudflare: stateful Durable Objects where the agent identity <em>is</em> the addressable compute unit); Railway, which <a href="https://www.prnewswire.com/news-releases/railway-raises-100-million-series-b-as-ai-pushes-todays-cloud-infrastructure-past-its-limits-302667768.html">raised $100 million in January</a> explicitly for this and whose founder Jake Cooper replied that there&#8217;s &#8220;a prize on offer worth playing for&#8221;; <a href="https://exe.dev">exe.dev</a> on the run-phase auth axis; and then agentuity, <a href="https://islo.dev">islo.dev</a>, <a href="https://hostess.sh">hostess.sh</a>, <a href="https://cnap.tech">cnap.tech</a>, and a long tail of less-evaluable seed-stage projects.</p><p>This spec serve as the obvious set of requirements that follows from the deployment shape, legible to anyone who has spent thirty minutes operating an agent against real infrastructure. The reason a dozen vendors can simultaneously claim &#8220;we&#8217;re working on this&#8221; is that the spec is not a secret. What&#8217;s hard is shipping it.</p><p>AWS won&#8217;t win this race, and it&#8217;s not because AWS doesn&#8217;t understand the requirements. It&#8217;s because the org structure can&#8217;t ship them. Twenty years of accumulated surface area, three thousand product managers with stakes in keeping their service distinct, and a billing system designed in 2008 to make it hard to comparison-shop against itself are not fixable from inside AWS. They&#8217;re only fixable by starting from a clean slate, which is what the clean-slate vendors are doing.</p><h2>The Heroku question</h2><p>Even if a vendor ships the entire spec, do they still lose? The Heroku pattern is the obvious template: great for proofs of concept, fine at modest scale, but the moment the business takes off, someone dispatches Claude Code to migrate the workload to AWS because that&#8217;s where the enterprise procurement, the compliance surface area, and the volume discounts live. Replace &#8220;place where companies start&#8221; with &#8220;place where indie devs run their agents&#8221; and you have a real risk to the entire thesis.</p><p>Here&#8217;s why it might not repeat. Heroku&#8217;s value proposition was developer experience: <code>git push deploy</code>, automatic Postgres, the Procfile abstraction. Those are workflow primitives, and AWS replicated them well enough to drain the at-scale customer. Amplify, App Runner (deprecated at the end of April, RIP), Lightsail, and Elastic Beanstalk are all &#8220;Heroku, but it&#8217;s on AWS so your CFO is happy.&#8221; None of them excellent; all of them sufficient.</p><p>The agent-native cloud&#8217;s value proposition is operational, not workflow-oriented. The operational primitives: capability-bounded sessions, time-travel rollback, hard budget caps with first-class agent identity attached. These are architectural primitives that would require AWS to refactor IAM, CloudTrail, and the billing system simultaneously. That&#8217;s not a console UI ship but a five-year coordinated rebuild across organizations that have spent twenty years optimizing for incompatible goals. An enterprise that has built operational practice around &#8220;my agents have first-class identities and hard budget caps&#8221; is migrating <em>away</em> from safety to go to AWS, not toward better economics. That&#8217;s a different migration vector than Heroku faced.</p><p>The risk is that AWS doesn&#8217;t have to ship the primitives well; they only have to ship something procurement will accept as &#8220;good enough&#8221; alongside existing AWS spend. The bar for keeping an enterprise customer isn&#8217;t &#8220;match Vercel on agent safety,&#8221; it&#8217;s &#8220;give the CFO a story they can tell the board about consolidating on one vendor.&#8221; <a href="https://aws.amazon.com/bedrock/guardrails/">Bedrock Guardrails</a> today doesn&#8217;t clear half the spec: it&#8217;s content filtering but not capability-bounded sessions or first-class agent identity. But &#8220;doesn&#8217;t pass the spec&#8221; and &#8220;good enough to win the renewal&#8221; are different bars, and AWS only has to clear the second.</p><p>The realistic call is that the agent-native cloud may end up serving two distinct populations: indie developers and small teams where the platform is the value, and enterprise pilots that eventually migrate to AWS once the workload matters enough that the CFO gets involved. My bet is that operational practice defined at the indie tier propagates up to where the enterprise workloads actually live, because that&#8217;s historically how new infrastructure categories have worked. But &#8220;eventually&#8221; is doing a lot of obnoxiously heavy lifting in that sentence.</p><h2>What I&#8217;ll be watching for</h2><p>The fourteen items above are the test. Hit them all, and I&#8217;ll concede you&#8217;ve built an agent-native cloud. Miss them, and you&#8217;ve built a marketing page with the word &#8220;agentic&#8221; in the headline; you can guess what my opinion is gonna be on that.</p><p>Vercel has a CEO willing to make a specific public bet on a thread written by a guy whose entire professional brand is calling out bullshit cloud claims. That deserves credit. Cloudflare&#8217;s response was different: the engineer who <a href="https://blog.cloudflare.com/agents-stripe-projects/">already shipped the most direct attempt at first-class agent identity</a> committed to shipping more of it. That also deserves credit, and now it&#8217;s a horse race.</p><p>Because neither company has yet shipped a credible answer to the hardest items: blast-radius primitives, capability-bounded sessions, or the cost preview API that would break fifteen years of AWS pricing opacity. Whether they get there before Railway, <a href="https://exe.dev">exe.dev</a>, a startup we haven&#8217;t heard of yet, or AWS shipping something passable enough to keep procurement happy is the question that determines who defines operational practice for agent infrastructure over the next several years, even if it doesn&#8217;t determine where every at-scale workload eventually runs.</p><p>That&#8217;s the bet. I&#8217;ll be watching the changelogs.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Subscribe for free, because the AI industry is hiring more PR firms than the unsubscribe button can keep up with.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence #2: The week AI labs became Palantir]]></title><description><![CDATA[Anthropic and OpenAI both stood up consulting arms this month. GitHub Copilot quietly admitted its subscription pricing never made sense. Vercel published the receipts.]]></description><link>https://artificialconfidence.com/p/artificial-confidence-2-the-week</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-2-the-week</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Tue, 19 May 2026 17:41:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!O-bH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There&#8217;s been a shift: the model layer used to be the prize, yet this week the AI industry quietly conceded there&#8217;s a great chance that it might be the loss leader. Anthropic launched a $1.5 billion enterprise services joint venture with Blackstone, Hellman &amp; Friedman, and Goldman Sachs on May 4. Seven days later, OpenAI countered with a $4 billion subsidiary called DeployCo, valued internally at $14 billion, anchored by 150 forward-deployed engineers (this cycle&#8217;s Hot New Job Title for &#8220;engineers who don&#8217;t embarrass themselves trying to hold a conversation&#8221;) acquired from a London consultancy you have not previously heard of. Three management consulting firms, Bain &amp; Company, Capgemini, and McKinsey, <a href="https://openai.com/index/openai-launches-the-deployment-company/">wrote checks</a> into the entity that seems explicitly designed to replace them. Either they&#8217;re hedging against their own obsolescence, or they would simply like to stop doing PowerPoints; both readings are supported by the press release. I too would like to stop doing PowerPoints, but that&#8217;s the job sometimes.</p><p>The week&#8217;s other money headlines line up underneath this thesis. If I can stuff them into one sentence like a flailing Labrador into an apartment bathtub: Cerebras traded, Anthropic floated $900 billion, OpenAI capped Microsoft&#8217;s revenue share at $38 billion, GitHub Copilot conceded its subscription pricing was never going to survive contact with agentic workloads, and Vercel published seven months of production data showing exactly why. That&#8217;s great, but to see most of what actually shipped, you had to read past the IPO frenzy to find. And that, dear reader, is why I&#8217;m writing this.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>What Actually Changed (Adjusted For Spin)</h2><h3>Anthropic launched Claude Platform on AWS, demoting Bedrock to &#8220;legacy&#8221; in its own docs</h3><p>On May 11, Anthropic announced <a href="https://aws.amazon.com/blogs/machine-learning/introducing-claude-platform-on-aws-anthropics-native-platform-through-your-aws-account/">general availability</a> of Claude Platform on AWS, a direct Anthropic API surface accessed through an AWS account. Authentication is IAM. Billing is through AWS Marketplace. Audit goes through CloudTrail. AWS is reduced to authentication, audit, and billing, which they do well&#8212;but none have anything to do with running the actual product. The hyperscaler has been demoted to &#8220;being Stripe, if Stripe&#8217;s UX was absolute dogshit.&#8221; Note: the inference runs on Anthropic-managed infrastructure outside the AWS security boundary, which is the sort of sentence that would have made a compliance officer eat a stapler in 2024 and is now a marketing bullet that the compliance officer is debating eating instead.</p><p>The framing AWS chose is &#8220;additive.&#8221; The framing Anthropic&#8217;s own docs chose is less beneficial to AWS, and consequently more honest. The existing Bedrock integration is now <a href="https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy">labeled</a> &#8220;legacy Amazon Bedrock integration.&#8221; Claude Opus 4.7, Anthropic&#8217;s current flagship, <a href="https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy">does not have an ARN-versioned model ID</a> on the legacy interface at all, which is the corporate equivalent of moving someone into a smaller office next to the printer spewing toner into the air and forgetting to give them the new keycard.</p><p>I clock two details the press release did not lead with. First, <a href="https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws">migrating from Bedrock</a> changes the SigV4 signing context, the base URL, the API format, the model IDs, the SDK client, the streaming format, the request headers, and the region availability, with an implicit customer message of &#8220;good luck, asspony.&#8221; Eight independent changes is not a &#8220;migration,&#8221; it&#8217;s a goddamned rewrite. Second, <a href="https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws">negotiated discounts and AWS Marketplace private offers do not transfer automatically between Bedrock and Claude Platform on AWS</a>. Translation: if you spent a quarter negotiating your Bedrock pricing, you get to spend another quarter negotiating again, and your existing EDP commit is not portable. For added fun, the <a href="https://docs.aws.amazon.com/claude-platform/latest/userguide/billing.html">AWS-side pricing structure</a> on this brings new meaning to &#8220;absurd:&#8221;</p><p><em>Usage is denominated in Claude Consumption Units (CCUs) at $0.01 USD per CCU. The CCU price is fixed and never discounted. Anthropic rates your token usage in USD at standard per-model, per-feature rates, applies any negotiated discount, then converts the result to CCUs at $0.01 per CCU. Discounts result in fewer CCUs metered, not a lower CCU price. CCUs are not prepaid credits; there is no CCU balance or commitment.</em></p><p>That&#8217;s a lot of words to say &#8220;you will not know what any of this costs until the bill shows up.&#8221;</p><h3>GitHub Copilot is moving to usage-based billing June 1, and the multipliers portend darkness</h3><p>GitHub <a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/">announced</a> that on June 1, all Copilot plans transition from premium request units to GitHub AI Credits, where one AI Credit equals one cent. Token consumption gets billed at published API rates, base subscription prices stay the same, code completions stay unlimited, and everything agentic gets metered.</p><p>The honest version of why came from GitHub&#8217;s own chief product officer, Mario Rodriguez: <a href="https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/">&#8220;a quick chat question and a multi-hour autonomous coding session can cost the user the same amount.&#8221;</a> It&#8217;s the kind of sentence you give a Senate committee when you have stopped trying to pretend the previous answer made sense; you bury the implementation in complexity that the senator from Pennsyltucky has no hope of answering, since they used to be a surgeon instead of a cloud economist.</p><p>It still doesn&#8217;t make a whole ton of sense as to why, so we go deeper to the most useful disclosure: the <a href="https://docs.github.com/en/copilot/reference/copilot-billing/model-multipliers-for-annual-plans">annual-plan model multiplier table</a>. For annual subscribers who stay on premium request billing after June 1, the Claude Opus 4.7 multiplier moves from 7.5x to 27x. The GPT-5.4 multiplier moves from 1x to 6x. GPT-4.1, previously a free 0x model under paid plans, is being pulled from the free tier entirely. The &#8220;prices aren&#8217;t changing&#8221; framing is technically accurate and also wildly misleading. The math under the prices changed by between four and six times depending on which frontier model you use. This is GitHub negotiating a contract cancellation through pricing. Annual subscribers running Opus 4.7 in agentic mode are being told, in the language of multipliers, that their existing contract no longer makes economic sense for GitHub, and would they please consider switching to monthly. Anyone who has ever taken a corporate &#8220;voluntary&#8221; buyout will recognize the structure.</p><p>If your team was paying $10 a month for Copilot Pro and burning Opus 4.7 in agentic mode, the unit economics you have been depending on were never real. The bill that arrives in July is the first one that reflects what your usage actually costs to serve, and a whole lot of folks are very much not going to enjoy this experience.</p><h3>OpenAI deleted DALL&#183;E 2 and DALL&#183;E 3 from the API on May 12</h3><p>Not deprecated like an AWS deprecation, deprecated like a Google deprecation: completely removed. The DALL&#183;E 2 and DALL&#183;E 3 model snapshots are <a href="https://developers.openai.com/api/docs/changelog">no longer available through the OpenAI API</a> as of May 12. The Realtime API Beta got the same &#8220;Old Yeller&#8221; treatment the same day. Your migration paths are gpt-image-2, gpt-image-1, gpt-image-1-mini, and the GA Realtime API respectively. Honestly, the hardest part of modern AI is teasing meaning from the model strings. Christ, I never thought I&#8217;d be nostalgic for AWS&#8217;s crap-ass naming &#8220;strategy.&#8221; Sure, Amazon DocumentDB (with MongoDB Compatibility) is a bad name, but at least you knew what the hell it was for.</p><p>If you&#8217;re one of the developers using the public OpenAI Image API in 2024, and you were not paying attention to the deprecation calendar, congratulations: you shipped a broken product last Tuesday. I am old; one of the things I always appreciated about AWS is that it&#8217;s vanishingly rare where their deprecations mean a thing that worked last week is broken this week. Meanwhile, that&#8217;s kinda the lived experience of being a Google customer. You get used to rapid change, invariably by surprise.</p><h3>IBM announced Red Hat AI Inference on IBM Cloud, GA May 22</h3><p>On the other end of the change continuum, over in IBM-land they&#8217;re launching a serverless inference API too. <a href="https://newsroom.ibm.com/2026-05-12-ibm-announces-red-hat-ai-inference-and-red-hat-openShift-virtualization-service-on-ibm-cloud">Powered by vLLM</a>, OpenAI-compatible API, the catalog includes Granite, Mistral-Small-3.2, Llama 3.3 70B Instruct, GPT-OSS-120B, and Nemotron-3-Nano-30B-FP8... If that doesn&#8217;t sound painful enough, it&#8217;s billed through IBM Cloud IAM. This is Bedrock and Vertex and AI Foundry, only with an IBM logo on it. Every hyperscaler and also IBM now sells inference-with-IAM as the product. The audit trail, business relationships, and significant install base comprising various forms of hostagetaking is the new moat.</p><div><hr></div><h2>Follow The Money (Or Watch It Follow Itself)</h2><h3>Cerebras traded, priced at 111x trailing revenue, then dropped 10%</h3><p><a href="https://www.cerebras.ai/press-release/cerebras-systems-announces-pricing-of-initial-public-offering">Cerebras Systems priced</a> at $185 on May 13, sold 30 million Class A shares, and raised $5.55 billion. Shares opened at $350 on May 14, <a href="https://techcrunch.com/2026/05/14/cerebras-raises-5-5b-kicking-off-2026s-ipo-season-with-a-bang/">peaked at $385</a>, closed the first day at $311.07, and dropped 10% on Friday. The marketed range moved from $115&#8211;125 to $150&#8211;160 to $185 in the days before pricing, which is what happens during a roadshow when you can feel and also unfortunately smell the room. At the IPO price, the implied fully diluted valuation was $56.4 billion. At the day-one peak, it was <a href="https://www.thestreet.com/investing/stocks/cerebras-stock-faces-sharp-reality-check-after-massive-5-5b-ipo-debut-drops-10">north of $120 billion</a>.</p><p>You&#8217;ll have to forgive me, but I&#8217;m from the 1900s, an era where it seems money meant something different than it does today. Cerebras&#8217;s FY25 revenue was <a href="https://www.investing.com/analysis/cerebras-48-billion-ipo-tests-the-markets-inference-bet-200680080">$510 million</a>, up 76% from $290 million the year before. At the IPO price, that valued the company at roughly 111 times trailing revenue. At the day-one close, more than 180 times. The Cerebras pitch is that this is not a normal chip company being valued like a chip company, it is scarce AI infrastructure being valued like scarce AI infrastructure, and the difference is the entire bull case. <strong>The Cerebras bull case is that we will, in fact, never have enough compute. The Cerebras bear case is that we will. Both cases were priced at $185 a share.</strong></p><h3>Vercel published April production data, and the labs are not competing on the same axis</h3><p>On May 12, Vercel <a href="https://vercel.com/blog/ai-gateway-production-index">published</a> seven months of AI Gateway production data covering more than 200,000 unique teams. The headline numbers for April: Anthropic took 61% of spend on 26% of token volume. Google took 21% of spend on 38% of volume. OpenAI took 12% of spend on 13% of volume, with spend share roughly tripling between March and April after the GPT-5.4 and GPT-5.5 releases. The labs are not competing for the same call. Anthropic is winning the high-stakes layer. Google is winning the high-volume-low-cost layer, which was basically the only pitch that Amazon&#8217;s Nova models had. And OpenAI is winning whatever it just shipped last week. Vercel&#8217;s own framing is that <a href="https://vercel.com/blog/ai-gateway-production-index">&#8220;spend follows the cost of being wrong&#8221;</a>, which is the cleanest one-line summary of inference economics anyone has shipped this year.</p><p>The same dataset shows 22.2% of AI Gateway requests in April ended with a tool call, carrying 58.9% of total token volume. The agentic share roughly doubled from October. The cost surface of production AI is now shaped like an agent, not a chat, and at the top of the request-volume curve, the average team is routing across <a href="https://vercel.com/blog/ai-gateway-production-index">35 distinct models</a>. The standard story about lab lock-in inverts the higher you go on the curve. Lab lock-in is a sales pitch. Routing graphs are infrastructure.</p><p>Vercel has skin in this game; the AI Gateway is their pitch to be the routing layer between those workloads and the labs, and &#8220;look at the multi-model fleets at the top of the curve&#8221; is exactly the argument a company selling routing infrastructure would want to make. The data is still the cleanest production-traffic read anyone has published, in part because nobody else with the volume has been editorially willing. I&#8217;m hesitant to overindex on their list of top providers, because unless customers override it the default selection is &#8220;whatever Vercel wants to use.&#8221; One wonders if they&#8217;re making these decisions based on their own commercial terms with various inference providers.</p><h3>Anthropic&#8217;s valuation cycle is now playing speed chess</h3><p>In February, Anthropic closed its Series G at a <a href="https://www.saastr.com/anthropic-just-hit-14-billion-in-arr-up-from-1-billion-just-14-months-ago/">$380 billion post-money valuation</a> on a $30 billion raise led by GIC and Coatue (gesundheit). On April 29, <a href="https://techcrunch.com/2026/04/29/sources-anthropic-could-raise-a-new-50b-round-at-a-valuation-of-900b/">TechCrunch reported</a> Anthropic had received preemptive offers at $800&#8211;900 billion and was sizing a $40&#8211;50 billion round. On May 12, <a href="https://www.bloomberg.com/news/articles/2026-05-12/anthropic-in-talks-to-raise-30-billion-at-900-billion-valuation">Bloomberg reported</a> the talks had crystallized into &#8220;at least $30 billion&#8221; at &#8220;more than $900 billion&#8221; pre-money, closing by the end of this month.</p><p>Four numbers, three weeks, one company. What&#8217;s interesting is that each leak walked the prior number in a particular direction. The raise size started at $40&#8211;50 billion (April 29), then narrowed to $30 billion (May 12). The valuation floor went from $800 billion (April 29) to $900 billion (May 12). The pattern is what it looks like: a series of trial balloons sized to discover the elastic limit of the room. Compare to the same company at <a href="https://venturebeat.com/technology/anthropic-says-it-hit-a-30-billion-revenue-run-rate-after-crazy-80x-growth">$61.5 billion in March 2025</a> and you have a roughly fifteen-fold private-market valuation move in fourteen months, which is statistically indistinguishable from a meme stock with a science publication.</p><p>The (Reported! We have nothing concrete!) revenue figure has done its own dance. End of 2025: <a href="https://www.understandingai.org/p/it-still-doesnt-look-like-theres">$9 billion</a>. Mid-February: $14 billion. Late February: $19 billion. April 7: <a href="https://venturebeat.com/technology/anthropic-says-it-hit-a-30-billion-revenue-run-rate-after-crazy-80x-growth">$30 billion</a>, per CFO Krishna Rao. April 29: <a href="https://techcrunch.com/2026/04/29/sources-anthropic-could-raise-a-new-50b-round-at-a-valuation-of-900b/">TechCrunch sources said</a> &#8220;closer to $40 billion.&#8221; Some mid-May reports cite $44 billion. OpenAI, which has its own reasons to argue this, <a href="https://venturebeat.com/technology/anthropic-says-it-hit-a-30-billion-revenue-run-rate-after-crazy-80x-growth">maintains</a> the $30 billion figure is overstated by approximately $8 billion on a gross-versus-net cloud revenue accounting argument, which would make the comparable number $22 billion. If you would like to know Anthropic&#8217;s annualized revenue today, please specify a date, a source, and which Magic 8-ball you consulted. They will not agree, and neither will Anthropic.</p><p>The growth itself is clearly real; I mean, eight of the Fortune 10 are paying customers (who the hell are the two holdouts, and can we talk?). One thousand of their customers spend over $1 million a year (theoretically on purpose), doubled from the February disclosure. Claude Code reached <a href="https://www.saastr.com/anthropic-just-hit-14-billion-in-arr-up-from-1-billion-just-14-months-ago/">$2.5 billion in run-rate revenue</a> within nine months of public launch. I want to be clear here: I&#8217;m not an idiot who denies reality, and I don&#8217;t have an agenda. My skepticism isn&#8217;t around whether the revenue exists, but rather which number, on which day, with which accounting treatment, ends up in the S-1.</p><h3>OpenAI capped Microsoft&#8217;s revenue share at $38B, $54B below Microsoft&#8217;s planning target</h3><p><a href="https://bilyonaryo.com/2026/05/13/openai-microsoft-agree-to-cap-revenue-sharing-at-38-billion/technology/">The Information reported</a> on May 11 that OpenAI and Microsoft agreed to cap total revenue-sharing payments at $38 billion, which is coincidentally how much the first publicly announced <a href="https://www.aboutamazon.com/news/aws/aws-open-ai-workloads-compute-infrastructure">OpenAI AWS deal was worth</a> in November, so maybe that&#8217;s the default amount in OpenAI&#8217;s QuickBooks installation or something. Microsoft had <a href="https://techresearchonline.com/news/openai-microsoft-revenue-sharing-deal/">internally been targeting</a> approximately $92 billion in returns from its OpenAI stake, per planning documents disclosed in the Musk v. Altman trial. The cap takes roughly $54 billion off Microsoft&#8217;s modeling and puts it back on OpenAI&#8217;s side of the table, which is exactly the kind of number you want to wave at IPO bookbuilders. Meanwhile Microsoft will presumably make up the shortfall and then some by putting ads into the GitHub service outage notifications.</p><p>In the same renegotiation, Microsoft&#8217;s license to OpenAI models was extended to 2032 but also made non-exclusive. OpenAI can now <a href="https://www.businesstoday.in/technology/story/openai-microsoft-cap-ai-revenue-sharing-payouts-at-38-billion-530973-2026-05-12">serve all its products across any cloud provider</a>. Microsoft&#8217;s previous revenue share to OpenAI was eliminated, leaving the cash flow one-directional. Read together, this is the contractual end of OpenAI&#8217;s Azure-exclusive era, made just visible enough that a public-market investor reading the S-1 will not accidentally believe the words &#8220;strategic partnership&#8221; mean anything specific.</p><h3>OpenAI launched DeployCo, raised $4B, and bought 150 Palantir-style engineers</h3><p>OpenAI&#8217;s corporate ADHD struck again as they <a href="https://openai.com/index/openai-launches-the-deployment-company/">announced DeployCo</a> on May 11, a majority-owned subsidiary capitalized with <a href="https://www.ciodive.com/news/openai-deployment-company-4-billion-ai-consulting-integration/819942/">more than $4 billion</a> from nineteen investors led by TPG. The implied valuation reported by Axios is <a href="https://letsdatascience.com/blog/openai-deployment-company-4b-tpg-tomoro-may-11-2026">$14 billion</a>, which is the number you produce by assuming a consulting practice that has existed for one day will scale faster than every consulting practice that has ever existed in the history of the world. The investor structure reportedly includes a 17.5% guaranteed return, which is the rate at which OpenAI has chosen to borrow $4 billion while calling the borrowing equity. That&#8217;s a similar guaranteed rate of return to that of many crypto emails lurking in my spam folder from 2019.</p><p>Three management consulting firms wrote checks: Bain &amp; Company, Capgemini, and McKinsey. My snark aside, companies are generally not run by idiots. Therefore, the polite reading is they are buying option value on the disruption of their own business. The less polite but spot-on reading is they have correctly priced the future cost of saying no.</p><p>Concurrent with the launch, OpenAI agreed to acquire <a href="https://thenextweb.com/news/tomoro-openai-deployment-company-consulting">Tomoro</a>, an Edinburgh-and-London consultancy founded in 2023 in alliance with OpenAI, employing approximately 150 forward-deployed engineers. At a $14 billion unit valuation, those engineers are valued at approximately $93 million per head, which is generous even by 2026 AI hiring standards. Anthropic shipped the same play seven days earlier with <a href="https://www.cnbc.com/2026/05/04/anthropic-goldman-blackstone-ai-venture.html">Blackstone, Hellman &amp; Friedman, and Goldman Sachs</a> on a $1.5 billion joint venture. Both labs have now formally conceded that the company that sells the model is not necessarily the company that captures the margin on its deployment; these are likely the early days of the frontier labs devouring their own ecosystems as pressure to show revenue builds.</p><p>The deeper reason I suspect drives these moves is that token revenue has structural unit economics problems that a public market analyst will, sooner rather than later, notice. Consultant time does not. Forward-deployed engineering is a revenue line that gets booked in dollars, not in inference losses, and converts cleanly to a chart that ends with the line going up. The labs aren&#8217;t pivoting to consulting because consulting is a great business. They&#8217;re pivoting to consulting because consulting is the only revenue line on the deck that does not require a footnote.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!O-bH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!O-bH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png 424w, https://substackcdn.com/image/fetch/$s_!O-bH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png 848w, https://substackcdn.com/image/fetch/$s_!O-bH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png 1272w, https://substackcdn.com/image/fetch/$s_!O-bH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!O-bH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated image" title="Generated image" srcset="https://substackcdn.com/image/fetch/$s_!O-bH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png 424w, https://substackcdn.com/image/fetch/$s_!O-bH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png 848w, https://substackcdn.com/image/fetch/$s_!O-bH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png 1272w, https://substackcdn.com/image/fetch/$s_!O-bH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cabdfe0-d9b0-4580-b03b-76fdf5fc39e7_5504x3072.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Reliability: A Brief Retrospective</h2><p>The Claude <a href="https://status.claude.com/">status page</a> records at least one investigated incident on May 12, 13, 14, 15, 16, and 18. The AI fanboys will no doubt point out that taking Sunday off has biblical precedent, and we&#8217;re closer than ever to summoning God via JSON. The May 13 cluster includes two separate investigations totaling roughly two and a half hours. The May 14 investigation lasted about two hours. The May 15 incident is the most interesting one editorially: the status update specifically <a href="https://status.claude.com/">notes</a> that &#8220;success rates for Opus 4.7 have returned to normal&#8221; while Opus 4.6 and Sonnet 4.6 were still degraded. The newer flagship recovered first. The older model and the smaller cheaper one stayed down longer.</p><p>The 90-day uptime numbers on the same status page tell a similar story by tier. Claude API: 98.99%. Claude Code: 99.14%. Claude Cowork: 99.45%. Claude for Government: 99.87%. The government tier gets approximately eight times less downtime than the public API, which is a useful way to think about exactly how much your federal contracting line item should cost to be worth it. OpenAI&#8217;s <a href="https://status.openai.com/">equivalent number</a> is 99.82%.</p><p>The number that offers more insight than either of those is the one Vercel <a href="https://vercel.com/blog/ai-gateway-production-index">published</a> on May 12 from seven months of AI Gateway data. About 3.5% of requests on the gateway end up rescued by failover to a healthy alternative. Measured by tokens, the rescue rate runs at 5.1%. Measured by dollars, 4.9%. The expensive end of the workload, long contexts, multi-step agent runs, heavy reasoning calls, is also the end most likely to need rescuing. A provider&#8217;s SLA measures request-level uptime. A production application experiences cost-weighted uptime, and the two blow themselves apart on exactly the calls that paid for the model. <strong>If your CFO is reading the SLAs and believing them, the CFO is reading the wrong document.</strong></p><div><hr></div><h2>The Hype Audit Department</h2><p>It&#8217;s worth saying out loud, because the prospectus does not. Cerebras shipped 30 million shares to public markets last week on the strength of a 76% revenue growth rate from $290 million to $510 million. That growth is real, or at least &#8220;real enough that if it&#8217;s not somebody will theoretically be going to jail.&#8221; The customers driving it are two government-funded UAE entities operating in the same emirate under the same sovereign sponsorship, plus a Master Relationship Agreement with OpenAI whose payments do not start materializing in the income statement until 2027 and whose existence assumes that OpenAI will be a buyer of physical inference compute four years from now in the volumes its current cap-table mathematics requires.</p><p>The phrase &#8220;diversified customer base&#8221; appears in the prospectus. The phrase &#8220;the same emirate&#8217;s two largest AI procurement vehicles&#8221; does not. The first phrase is technically accurate. The second is also technically accurate, and if we&#8217;re being direct it&#8217;s the one that should be priced in. Cerebras&#8217;s 86% two-customer concentration in 2025 is one percentage point higher than its 85% single-customer concentration in 2024. The change is a new LLC name on the second-largest line, not a new geography.</p><p>Both LLCs sit inside the same sovereign portfolio, and the prospectus knows it. On <a href="https://www.sec.gov/Archives/edgar/data/2021728/000162828026035214/cerebras-424b4.htm">page 22</a>, the company discloses that &#8220;G42 and MBZUAI are considered related parties with respect to each other as defined by Accounting Standards Codification 850.&#8221; For those of you who aren&#8217;t giant nerds, ASC 850 is the accounting rule that requires companies to flag related-party connections; Cerebras checked the box, used the accounting-standard citation as the entire structural acknowledgment, and stopped, hoping everyone else would too. The prospectus never specifies the relationship.</p><p>The relationship is this: G42 is chaired and controlled by Sheikh Tahnoun bin Zayed Al Nahyan, the UAE&#8217;s National Security Advisor since 2016, who oversees roughly $1.5 trillion in sovereign capital and was deputy national security advisor at the time of Project Raven, <a href="https://www.reuters.com/investigates/special-report/usa-spying-raven/">the surveillance program Reuters documented in 2019</a> that hired former NSA personnel to spy on American citizens, journalists, and dissidents. MBZUAI (pronounced like an Amazon seller who&#8217;s about to scam you) stands for &#8220;Mohamed bin Zayed University of Artificial Intelligence&#8221; and is named after his brother, Sheikh Mohamed bin Zayed Al Nahyan, the President of the UAE. E.</p><p>The closest the S-1 comes to acknowledging any of this is one passing reference, in the same risk factor, to &#8220;laws or regulations applicable to OpenAI, G42 or MBZUAI, or the United Arab Emirates.&#8221; That is the entire UAE disclosure. The investor is left to assemble the rest, which is the work AC exists to do.</p><p>The bear case for CBRS is not &#8220;the growth is fake,&#8221; it&#8217;s that &#8220;Mohamed bin Zayed University of Artificial Intelligence&#8221; and &#8220;G42&#8221; are not the names of a diversification strategy any more than &#8220;we are diversified between a guy and also his brother, both of whom have diplomatic immunity&#8221; is.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EPom!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EPom!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png 424w, https://substackcdn.com/image/fetch/$s_!EPom!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png 848w, https://substackcdn.com/image/fetch/$s_!EPom!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png 1272w, https://substackcdn.com/image/fetch/$s_!EPom!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EPom!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Generated image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Generated image" title="Generated image" srcset="https://substackcdn.com/image/fetch/$s_!EPom!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png 424w, https://substackcdn.com/image/fetch/$s_!EPom!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png 848w, https://substackcdn.com/image/fetch/$s_!EPom!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png 1272w, https://substackcdn.com/image/fetch/$s_!EPom!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feca6c12a-b606-409a-9a23-e28fe8725faa_5504x3072.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>One last thing</h2><p>This week the AI industry decided what kind of company it actually is. It turns out that it&#8217;s not a model vendor with a consulting practice attached (usually called &#8220;Professional Services&#8221;). Rather, it&#8217;s a consulting practice with a model vendor attached, and the model vendor&#8217;s job is to keep the consulting practice differentiated from Accenture. Anthropic and OpenAI both put $5.5 billion of investor capital toward this thesis in the same fortnight. Cerebras went public on the back of a procurement relationship with two foreign-government-adjacent buyers. Microsoft accepted a $54 billion haircut on its OpenAI returns so the cap table would look right for an IPO. GitHub admitted the subscription pricing it has been running for two years was never going to survive the agentic workloads it explicitly built the product around. Vercel published the data. The model layer is still the thing investors are buying, but it is not the thing they are paying for.</p><p>If you run AI workloads and you have not renegotiated your cloud commit this quarter, your counterparty just made it harder for you. If you sell consulting and you have not noticed that the labs are now your competitors, your counterparty also just made it harder for you. And if you have a Copilot Pro subscription on auto-renew and you have not looked at the June multiplier table, you are about to be one of the case studies in a future issue of Artificial Confidence. Either way, the bill changed.</p><p>See you next week.</p><p>&#8212; C</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Artificial Confidence #1: AWS gave the agents a credit card]]></title><description><![CDATA[Inaugural issue. Microsoft Research benchmarked the agents and they're not ready; AWS gave them a credit card anyway.]]></description><link>https://artificialconfidence.com/p/artificial-confidence-1-aws-gave</link><guid isPermaLink="false">https://artificialconfidence.com/p/artificial-confidence-1-aws-gave</guid><dc:creator><![CDATA[Corey Quinn]]></dc:creator><pubDate>Tue, 12 May 2026 19:13:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!diRt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!diRt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!diRt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!diRt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!diRt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!diRt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!diRt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/adb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2862304,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://artificialconfidence.substack.com/i/197302431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!diRt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!diRt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!diRt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!diRt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadb78aee-cd90-45ad-8079-6fb1ecf1c334_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><p>Hello and surprise to many of you; welcome to the inaugural issue of &#8220;Artificial Confidence.&#8221; Here, I cover the AI news from roughly the past week that doesn&#8217;t quite fit into <a href="https://www.lastweekinaws.com">Last Week in AWS</a>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>&#8220;Last Week in AWS&#8221; came from something I desperately wanted: a source to round up the stuff from AWS&#8217;s cloud ecosystem that <em>mattered</em> to customers. We&#8217;re seeing a similar content spew in the AI space: lots of hype, lots of noise, yet remarkably low signal. I&#8217;ve grown weary of waiting for someone else to do it, so it&#8217;s time to be the change I want to see in the world. I want to bring an overheated, overhyped space to life in a way that humans actually care about without spending hours a day drudging through the muck. I want to surface the things that may have slipped past unremarked under a deluge of <a href="https://karlbode.com/ceo-said-a-thing-journalism/">CEO said a thing</a> style &#8220;journalism.&#8221; And I want to write it myself; the mortal sin of so much AI generated content nowadays is people believing that you&#8217;ll take the time to read something they couldn&#8217;t even be bothered to <em>write</em>.</p><p>If this isn&#8217;t for you, I understand completely; whack the unsubscribe link. Your &#8220;Last Week in AWS&#8221; subscription will remain unaffected; go ahead and cancel that too if you&#8217;re annoyed with me and were waiting for an excuse. I get it; even <strong>AWS</strong> doesn&#8217;t talk about AWS releases the way they once did. But I hope you&#8217;ll stick around.</p><div><hr></div><p>Vendor story this week: AI agents can autonomously do everything. Research story this week: no the hell they cannot. <a href="https://arxiv.org/abs/2604.15597">Microsoft&#8217;s own scientists</a> found the agents corrupt 25% of multi-step work on average, and that adding tools makes performance 6% <em>worse</em>. AWS, the same week, <a href="https://aws.amazon.com/blogs/machine-learning/agents-that-transact-introducing-amazon-bedrock-agentcore-payments-built-with-coinbase-and-stripe/">announced</a> you can now give those agents a wallet. I have spent a decade watching AWS announce capabilities that arrived years before the safety infrastructure to support them; it&#8217;s refreshing to see the AI industry compress that timeline into a single news cycle.</p><div><hr></div><h2>What Actually Changed (Adjusted For Spin)</h2><h3>Claude Opus 4.7 raised prices without raising prices</h3><p>Anthropic shipped <a href="https://platform.claude.com/docs/en/about-claude/models/whats-new-claude-4-7">Claude Opus 4.7</a> on April 16 with what they have, repeatedly, called &#8220;unchanged pricing.&#8221; Five dollars per million input tokens, twenty-five per million output, identical to Opus 4.6, 4.5, and 4.1. The pricing page has been the very model of consistency.</p><p>The tokenizer, however, has not. Opus 4.7&#8217;s new tokenizer is denser, which is the polite engineering phrasing for &#8220;turns the same English sentence into more tokens than the old one because we have hilariously overcommitted to buy every GPU on the planet and must pretend to be able to pay for them somehow.&#8221; <a href="https://platform.claude.com/docs/en/about-claude/models/whats-new-claude-4-7">Anthropic&#8217;s docs put the multiplier at 1.0x to 1.35x</a>, with the upper end showing up on code, structured data, and non-English text. The number is filed on the &#8220;what&#8217;s new&#8221; page rather than the pricing page. That&#8217;s the kind of editorial decision you make when you would prefer the number not appear where purchasing decisions get made. The same page recommends &#8220;updating your <code>max_tokens</code> parameters to give additional headroom,&#8221; which is the advice you give people about to use more tokens than they were planning to, while hoping they aren&#8217;t astute enough to figure that out. A <a href="https://medium.com/@dev_tips/the-ai-price-hike-that-never-showed-up-on-the-pricing-page-your-bill-went-up-27-anyway-48a61265f3f3">practitioner write-up on Medium</a> benchmarked a real workload at a 27% bump on identical prompts.</p><p>This is more elegant than charging more for the same number. It is also less honest.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mcxs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mcxs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!mcxs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!mcxs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!mcxs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mcxs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3118953,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://artificialconfidence.substack.com/i/197302431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mcxs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!mcxs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!mcxs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!mcxs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7060269-901d-4d1e-9f07-2cb6f6aef52e_2752x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h3>AWS Bedrock AgentCore now lets agents pay for things</h3><p><a href="https://aws.amazon.com/blogs/machine-learning/agents-that-transact-introducing-amazon-bedrock-agentcore-payments-built-with-coinbase-and-stripe/">Announced May 7.</a> AI agents can now autonomously pay for APIs, MCP servers, web content, and other agents, via Coinbase CDP wallets or Stripe Privy (was &#8220;Shitr&#8221; taken?) wallets, with what the announcement repeatedly describes as &#8220;session-level spend limits,&#8221; whatever the hell that&#8217;s supposed to mean.</p><p>First: At last, I finally get to solve my biggest pain point as a customer: not being able to pay for things without human supervision. Er&#8230; wat? Does anyone actually have a problem doing this?</p><p>Second: I have done some looking, and I have not yet found a clear, durable definition of what a &#8220;session&#8221; is. Single API call? Single agent invocation? Single user-facing transaction? AWS, with their characteristic forthrightness, has provided exactly enough specificity to ship the feature and exactly enough ambiguity to ship the blog post. It&#8217;s now technically possible for an agent to burn $40,000 overnight against a misconfigured spend limit, an outcome that has been moved from &#8220;theoretical concern&#8221; to &#8220;forthcoming case study.&#8221; The post-mortem write-up is already in my saved-drafts folder, dated approximately six months from today.</p><h3>Amazon Q Developer is being deprecated on an unusually short timeline</h3><p><a href="https://aws.amazon.com/blogs/devops/amazon-q-developer-end-of-support-announcement/">Announced April 30.</a> Q Developer IDE plugins and paid subscriptions, which until very recently AWS was attempting to shove down our throats with zeal and gusto, reach end-of-life on April 30, 2027. Twelve months from announcement to &#8220;gone,&#8221; which is by AWS standards, brisk. The usual cadence is longer, but then again the usual cadence is also aligned with customers who are knowingly using the product.</p><ul><li><p><strong>May 15, 2026</strong> (this Friday): no new signups, not that that was a problem.</p></li><li><p><strong>May 29, 2026</strong>: Opus 4.6 disappears from Q Developer Pro.</p></li><li><p>Opus 4.7, the current Anthropic flagship, is available exclusively on Kiro; the replacement product you also have no interest in using.</p></li></ul><p>If you run Q Developer Pro and you have been pretending Kiro is not a thing, AWS would like you to know that you have approximately two weeks before they begin to migrate you on their schedule.</p><h3>Both Anthropic and OpenAI are now on Bedrock</h3><p><a href="https://aws.amazon.com/about-aws/whats-new/2026/04/bedrock-openai-models-codex-managed-agents/">Announced April 28</a>. GPT-5.5 and GPT-5.4 on Bedrock in limited preview. Codex on Bedrock. Also &#8220;Amazon Bedrock Managed Agents, Powered by OpenAI&#8221; which is the longest AWS product name of 2026, and that is saying something.</p><p>For two years, Bedrock has been &#8220;Claude on AWS, plus a bunch of other rando models you won&#8217;t use on purpose.&#8221; It is now &#8220;either of the top two US AI labs on AWS.&#8221; Anthropic&#8217;s special-est-friend status has been quietly rezoned to &#8220;one of two preferred partners,&#8221; which is the corporate-relationship equivalent of being informed your spouse has decided to start dating again.</p><p>AgentCore in GovCloud (US-West) <a href="https://aws.amazon.com/about-aws/whats-new/2026/05/bedrock-agentcore-launch-aws-govcloud-us/">also went live May 5</a>. Government workloads can now run agents. We&#8217;ll revisit in six months when something interesting happens.</p><div><hr></div><h2>Agents Got Powerful This Week. They Also Got Worse.</h2><h3>Microsoft&#8217;s own scientists: agents corrupt 25% of your work, and tools make it worse</h3><p>Microsoft Research published a paper on Monday with the genuinely on-brand title <a href="https://arxiv.org/abs/2604.15597">&#8220;LLMs Corrupt Your Documents When You Delegate.&#8221;</a> Unlike statements from the non-Research parts of Microsoft, it is exactly the paper the title suggests.</p><p>They created a benchmark called DELEGATE-52 which appropriately tests flows across fifty-two professional domains. Because the devil lives inside your corporate process, they feature twenty-interaction multi-step workflows. This brings us to their findings, ordered by how irritating each is to the marketing departments of every major AI vendor:</p><ul><li><p>Frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT-5.4) lose, on average, <strong>25% of document content</strong> over twenty interactions. The all-model average is <strong>50%</strong>. You are reading those numbers correctly. If you hear screaming coming from the C-suite, so are they&#8212;and you just learned a valuable thing about the demographics of this publication.</p></li><li><p>Of fifty-two domains tested, <strong>exactly one</strong> met the bar of &#8220;ready for delegation&#8221; (&#8805;98% accuracy after twenty rounds). That domain, completely unsurprisingly, is Python programming. Every other domain: accounting, music notation, crystallography; the actual knowledge work you would actually delegate to an actual agent? Yeah, that&#8217;s a harsher number you&#8217;re really not gonna like.</p></li><li><p>&#8220;Catastrophic corruption&#8221; (&#8804;80% score) occurred in <strong>80%+ of model/domain combinations</strong>.</p></li><li><p>The errors don&#8217;t accumulate gradually. They arrive in single 10-to-30-point drops. That&#8217;s the failure mode hardest to build an SLA against, and leads to questions that start with blistering profanity in the first sentence.</p></li><li><p>The somehow worse finding: <strong>adding an agentic harness with tools makes performance 6% </strong><em><strong>worse</strong></em><strong> on average.</strong> The entire architectural premise of the agentic movement currently being rammed down our throats (&#8220;give the model tools and it becomes more capable&#8221;) provably degrades outcomes in this benchmark.</p></li></ul><p>This is not &#8220;AI is useless.&#8221; It is narrower and more devastating: the exact vendor positioning that Anthropic, OpenAI, Microsoft, Google, and AWS are <em>all currently selling on the same Tuesday morning</em>, consisting of &#8220;hand the agent a multi-step task and walk away,&#8221; is empirically contradicted by Microsoft&#8217;s own employees. <a href="https://www.theregister.com/ai-ml/2026/05/11/microsoft-researchers-find-ai-models-and-agents-cant-handle-long-running-tasks/5238263">The Register</a> opens with &#8220;an intern who failed this much would be shown the door.&#8221; That is generous. An intern who lost 25% of a document gets a performance improvement plan. An intern who made the problem <em>worse</em> by adding a calculator gets introduced to a baseball bat after hours in some shops.</p><h3>The same week, AWS announced autonomous-spending agents</h3><p>Microsoft Research, Monday: agents corrupt a quarter of your work, and tools make that worse. AWS, the previous Thursday: now you can give those agents a wallet. I will leave these two stories next to each other and let you make the joke. Consider it a participatory newsletter.</p><h3>A British mathematician handed an agent a credit card</h3><p><a href="https://www.theregister.com/software/2026/05/05/british-mathematician-hands-openclaw-agent-a-credit-card/5228654">The Register, May 5.</a> An experimental run of an AI agent given payment authority that ended in password leaks, CAPTCHA chaos, and the kind of behavior you would expect from a sufficiently empowered toddler in a Best Buy. The headline is tabloid, but the experimental setup is approximately what AgentCore payments enables in production. The headline is also a more honest preview of where this leads than the AgentCore announcement is. It has to be.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6bwP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6bwP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!6bwP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!6bwP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!6bwP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6bwP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2372867,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://artificialconfidence.substack.com/i/197302431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6bwP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png 424w, https://substackcdn.com/image/fetch/$s_!6bwP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png 848w, https://substackcdn.com/image/fetch/$s_!6bwP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png 1272w, https://substackcdn.com/image/fetch/$s_!6bwP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c8a9cbb-a06c-495b-aa66-9f457ed55b5b_2048x2048.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><h2>Reliability: A Brief Retrospective</h2><p>Last week was unusually rough for AI infrastructure. Five days, four user-affecting incidents, four different vendors, and we&#8217;ll skip GitHub because at this point I don&#8217;t think they come here for the hunting anymore:</p><ul><li><p><strong>May 5</strong>: <a href="https://www.techtimes.com/articles/316343/20260505/google-gemini-down-now-ai-service-experiencing-outages-users-report-delays-errors-may-5.htm">Google Gemini degraded widely</a> starting around 8:44 AM EDT, free and paid both. Multimodal hit harder than text; that&#8217;s an instructive nugget about which paths apparently share infrastructure.</p></li><li><p><strong>May 7</strong>: <a href="https://www.theregister.com/off-prem/2026/05/07/ibm-cloud-evaporates-as-datacenter-loses-power/5234835">IBM Cloud lost power at a datacenter</a>. IBM &#8220;Cloud&#8221; is technically not an AI provider nor a real cloud, but enough AI runs on top of it that this counted.</p></li><li><p><strong>May 8</strong>: Significant Claude outage. The same afternoon, <a href="https://status.openai.com/history">OpenAI&#8217;s Responses API threw 404s for 35 minutes</a> after a bad deploy. Two of the three major US AI providers degraded within hours of each other.</p></li><li><p><strong>May 9</strong>: <a href="https://status.claude.com/">Claude Code on Web partial outage; Opus 4.1 elevated errors</a>.</p></li></ul><p><a href="https://isdown.app/status/anthropic">IsDown has logged 671 Anthropic incidents since June 2024</a>, showing incidents typically resolving within 246 minutes. Multiply that by 671 and you&#8217;re measuring uptime that&#8217;s comparable to your bank&#8217;s business hours. The industry&#8217;s response to this baseline is, apparently, to put autonomous spending capability on top of it on the theory that none of us is as dumb as all of us. If your foundation has four-hour outages every couple of weeks and your stated direction is &#8220;deploy agents that pay for things on top of that foundation,&#8221; I would respectfully suggest that your circus is missing one of their underperforming clowns. (<a href="https://dev.to/safdarali25/is-ai-down-today-full-status-report-for-chatgpt-claude-gemini-more-1i4g">Decent recap on DEV.to</a>.)</p><div><hr></div><h2>Follow The Money (Or Watch It Follow Itself)</h2><h3>Cerebras trades Thursday</h3><p>Cerebras (CBRS) hits Nasdaq Thursday morning, set to be approximately the seventh &#8220;biggest AI IPO of 2026 so far,&#8221; a title that has changed hands roughly every six weeks since January, in a year not yet half over. The price-range escalation <a href="https://www.cnbc.com/2026/05/11/cerebras-raises-ipo-range.html">moved through three acts in a fortnight</a>: $115-$125, then $125-$135, then $150-$160 on 20x oversubscription. That is the bankers&#8217; way of admitting they underpriced the offering so embarrassingly that they would, in retrospect, like a do-over with witnesses present.</p><p>At the top, Cerebras raises ~$4.8 billion at a $48.8 billion fully-diluted valuation, or approximately 96 times trailing revenue. <a href="https://www.sec.gov/Archives/edgar/data/2021728/000162828026025762/cerebras-sx1april2026.htm">Trailing revenue is $510 million for 2025</a>, with a reported <strong>47% net margin</strong>. The word &#8220;reported&#8221; is doing considerable structural work in that sentence. Headline GAAP net income: $237.8 million. Of which $363.3 million is a one-time non-cash gain from extinguishing a forward-contract liability tied to G42. Strip out the accounting and Cerebras <a href="https://futurumgroup.com/insights/cerebras-s-1-teardown-is-the-23b-wafer-scale-ipo-the-end-of-gpu-homogeneity/">posted a non-GAAP net loss of $75.7 million</a> for the year. Cerebras is &#8220;profitable&#8221; in roughly the same way you are profitable in a year you cleaned out the storage unit.</p><p>The prospectus is also unusually candid about customer concentration, in the way that suggests counsel concluded the SEC was going to ask anyway and did not want to be the ones holding the bag. Two UAE-based entities <a href="https://futurumgroup.com/insights/cerebras-s-1-teardown-is-the-23b-wafer-scale-ipo-the-end-of-gpu-homogeneity/">historically account for roughly 86% of revenue</a>, the kind of currently war-afflicted geographic dependence that turns &#8220;concentration risk&#8221; into &#8220;two phone calls and a passport.&#8221; They say that OpenAI&#8217;s $20+ billion 750-megawatt deal represents &#8220;a substantial portion of projected revenue over the next several years,&#8221; which is S-1 language for &#8220;if anything happens to that one phone call, the rest of this prospectus is fiction.&#8221; The other hyperscaler customer is AWS, which is, as everyone knows, deeply enthusiastic about taking hard dependencies on third-party compute and never abandons them halfway through.</p><p>Read alongside <a href="https://www.cnbc.com/2026/05/09/nvidia-embraces-ai-investor-topping-40-billion-in-equity-bets-2026.html">Nvidia&#8217;s $40+ billion in 2026 equity investments</a> (including, of course, a reported $30 billion stake in OpenAI), and the picture is this: Nvidia, the largest AI compute supplier, has placed its biggest equity bet on OpenAI; OpenAI is the largest customer of Cerebras; Cerebras is the AI compute supplier going public this week. This is what economists call &#8220;circular&#8221; and what regulators tend to call &#8220;some kind of obscene financial ouroboros we will be looking into in three years.&#8221;</p><h3>Snap and Perplexity quietly buried the $400M partnership</h3><p>The $400 million Snap-Perplexity deal (which was supposed to put Perplexity&#8217;s conversational search inside Snapchat for reasons the original announcement struggled to articulate clearly even at announcement time) has been <a href="https://techcrunch.com/2026/05/06/snap-says-its-400m-deal-with-perplexity-amicably-ended/">&#8220;amicably ended,&#8221;</a> per Snap&#8217;s Q1 earnings disclosure last week. &#8220;Amicably ended&#8221; is the financial-PR phrasing for &#8220;one or both parties walked into a meeting in February and could no longer remember what the slide deck was for.&#8221; Whatever the testing surfaced was bad enough that nine figures of pre-committed capital wasn&#8217;t enough to paper over it. In 2026, that is quaintly reassuring.</p><div><hr></div><h2>The Hype Audit Department</h2><h3>Mythos vs. cURL: one low-severity CVE, after all that</h3><p>Anthropic&#8217;s Mythos is the company&#8217;s flagship &#8220;too dangerous to release publicly&#8221; cybersecurity model, purportedly capable of identifying and exploiting security vulnerabilities at a level beyond what&#8217;s safe to put in general circulation. The marketing implication: Pandora&#8217;s box on legs. Project Glasswing, via the Linux Foundation, provides gated access to selected open-source maintainers so they can use it defensively. Daniel Stenberg, the cURL maintainer, was on the list.</p><p><a href="https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-vulnerability/">Mythos ran against the cURL codebase</a> and returned five &#8220;confirmed&#8221; vulnerabilities. Stenberg&#8217;s team reviewed them. Three were false positives pointing at things already documented in cURL&#8217;s own API docs. One was a non-security bug. The fifth&#8212;and only&#8212;actual vulnerability is a <strong>low-severity CVE</strong> shipping with cURL 8.21.0 in late June. In Stenberg&#8217;s words: &#8220;The flaw is not going to make anyone grasp for breath.&#8221; For context: AI tooling has contributed 200&#8211;300 bugfixes to cURL over the last 8&#8211;10 months, and Stenberg says modern AI analyzers are better than what came before. He just doesn&#8217;t think Mythos is meaningfully better than the other modern AI analyzers. He calls the hype &#8220;primarily marketing.&#8221; (<a href="https://www.theregister.com/security/2026/05/11/anthropics-bug-hunting-mythos-was-greatest-marketing-stunt-ever-says-curl-creator/5238111">The Register</a> hits harder than Stenberg actually did.)</p><p>The caveat: Stenberg never received hands-on access to Mythos. He signed up for Glasswing, but someone else ran the scan and sent him the report. He is, as of his Monday blog post, still waiting for direct access because the model is oh-so-scawy. The gating is tight enough that even the maintainers Anthropic is ostensibly trying to help cannot independently confirm or contest the danger claim. Combine that with <a href="https://www.theregister.com/software/2026/04/22/mythos-found-271-firefox-flaws-none-a-human-couldnt-spot/5223657">April&#8217;s Firefox audit</a> (271 flaws found, zero a competent human couldn&#8217;t have spotted) and a pattern emerges: every time someone qualified gets within evaluation distance of Mythos, they conclude it is unremarkable. Capable, but unremarkable. When the safety story IS the marketing story, you cannot tell them apart. I don&#8217;t think that&#8217;s an accident.</p><h3>Google: criminals already used AI-built zero-days in the wild</h3><p>On the same day The Register published the cURL piece, <a href="https://www.theregister.com/ai-ml/2026/05/11/google-says-criminals-used-ai-built-zero-day-in-planned-mass-hack-spree/5237982">Google&#8217;s Threat Intelligence Group reported</a> that criminals had already operationalized an AI-built zero-day in an attempted mass exploitation campaign. The defensive AI-vulnerability-finder, you will recall, is gated as too dangerous to release publicly. The offensive use is happening regardless. &#8220;Too dangerous to release&#8221; is the kind of framing that requires the bad guys to be waiting on the release schedule. They are, alas, not. Cynically, I wonder if the real reason not to release Mythos rhymes with &#8220;mompute schmortage.&#8221;</p><div><hr></div><h2>Where I&#8217;ll be</h2><p>The Duckbill team (y&#8217;know, my day job) has a busy May and June, and we&#8217;re using it as an excuse to host dinners at every stop. I&#8217;ll be at all of them, if that&#8217;s the kind of thing that influences your dinner plans.</p><p>First up: <a href="https://luma.com/pgbvditk">San Francisco on May 19th</a>, a small, off-the-record dinner about negotiating with hyperscalers. Jim Moses and I will be there, but this isn&#8217;t a presentation. It&#8217;s a conversation among people who&#8217;ve actually been in those rooms, with all the wit and sarcasm you&#8217;ve come to expect.</p><p>Then, because apparently we hate ourselves, we&#8217;re doing back-to-back AWS Summits in <a href="https://luma.com/fc51znep">LA on June 9th</a> and <a href="https://luma.com/8goovbdf">NYC on June 16th</a>, with dinners at both for people in cloud cost and FinOps who want to continue the conference conversation somewhere with better food.</p><p>Spots are limited and require approval. Not mine, of course; you&#8217;re all aces in my book.</p><div><hr></div><h2>One last thing</h2><p>If you read one thing this week, read <a href="https://arxiv.org/abs/2604.15597">the Microsoft paper</a>; do not let the agents spend your money unattended in the meantime.</p><p>See you next week.</p><p>&#8212; C</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://artificialconfidence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Artificial Confidence! Subscribe for free to receive new posts and see what I&#8217;m up to next.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>