Prince Mario-Max Schaumburg-Lippe: Google Antigravity Adds Claude 5.5 Coding Models

Google quietly did something last week that developers noticed immediately: it put a rival’s flagship models inside its own coding IDE. On October 3, Google added Anthropic’s Claude Opus 5.5 and Claude Sonnet 5.5 to the model selector in Antigravity, its agentic development workspace, for paying subscribers on the Google AI Pro and Google AI Ultra tiers.

The update closed a gap that had been sitting in plain sight. Antigravity’s model page had listed the new 5.5 models as unavailable since late September, and developers were asking when the current generation would show up. It showed up with no fanfare, no blog post, just two new entries in a dropdown. Which, honestly, might be the most Google way to ship anything.

What changed in the model lineup

Both additions are the reasoning “thinking” variants: Claude Opus 5.5 (thinking) and Claude Sonnet 5.5 (thinking). Access is gated to non-trial Google AI Pro and Google AI Ultra subscriptions, which run $19.99, $99.99 and $199.99 a month depending on the tier. Free accounts, the cheaper Plus tier, and Enterprise accounts don’t get either model, so the availability picture is narrower than a simple paid-versus-free split.

At the same time, Google set an expiration date on the old guard. Claude Opus 4.6, Claude Sonnet 4.6 and the open-weights GPT-OSS-120B are scheduled for removal on November 2. Gemini 3.1 Pro stays as the default, with Gemini 3.8, 3.7 and 3.6 Flash rounding out Google’s own options. If you are still running on 4.6 in Antigravity, you have about a month to move your workflows.

That retirement schedule matters more than it looks. When a platform kills a model version, every prompt, benchmark and test result built on it becomes history. Teams that treat model choice as casually as a dropdown setting will feel this one. The ones that pin versions and track which model produced which result won’t.

The models themselves are worth the slot

Claude Opus 5.5 shipped September 22 with a 20% price cut over Opus 5, dropping to $4 per million input tokens and $20 per million output. On benchmarks, Opus 5.5 posts 66.4% on Terminal-Bench 4.0, up from Opus 5’s 52.3%, and Anthropic says it matches the top score of OpenAI’s GPT-6 Astra on FrontierCode at roughly a fifth of the cost.

Sonnet 5.5, launched September 28, is the medium model built for everyday work: bug fixes, features with written specs, polished documents and slides. It runs more than 30% faster than Sonnet 5, costs up to 30% less for most work, and keeps Sonnet 5’s pricing at $2 per million input and $10 per million output tokens. One reported customer test completed a 680,000-line code migration in less than a day on Opus 5.5, which tells you where the frontier is on long-horizon agent work right now.

The real story is the bundling

Here’s what makes this update interesting beyond the version numbers. Google’s Antigravity has been model-agnostic from the start. It launched in November 2025 offering Claude Sonnet 4.5 and OpenAI’s GPT-OSS alongside Gemini, all billed through a Google account. That strategy just got extended to Anthropic’s current generation instead of the older models Google is now retiring.

Think about the math from a team’s perspective. Pay Anthropic directly and you are billed per token through its API. Pay Google a flat monthly subscription and you get Opus 5.5, Sonnet 5.5, Gemini 3.1 Pro and the open models under one invoice, with usage limits set by Google’s tier rather than Anthropic’s meter. For teams already inside Google Workspace or Google Cloud, there is an obvious gravitational pull to let Antigravity be the place where model comparison happens, instead of juggling three separate subscriptions.

And comparison is the operative word. Putting Anthropic’s models in the same dropdown as Gemini means every developer in Antigravity can run a live head-to-head every time they open a new session. That’s a confident move by Google. It says the company believes its orchestration layer, not any single model, is the product. A coding platform that locks a team into one model family builds a moat out of inconvenience. Google seems to have decided the better moat is the workspace itself.

What it means for working developers

The practical upshot is simple: your IDE is becoming the least bad place to answer the hardest question in AI coding right now, which is which model for which task. Opus 5.5 for the ambiguous, multi-file, security-sensitive work. Sonnet 5.5 for the well-specified everyday grind. Gemini for the massive codebases where a million-token context window pays rent. Having all three one click apart, on one bill, lowers the friction of picking right.

That matters because the coding-assistant market has spent 2026 fragmenting. Every model vendor wants you in its own environment, its own API, its own pricing scheme. The open-source inference world is moving the other way, with new platforms like Prime Intellect’s inference service letting teams serve frontier open models on their own GPUs. Antigravity sits in the middle: proprietary models, one subscription, no API keys to manage.

There’s also a quiet signal in what Google chose not to do. It could have kept the newest Anthropic models out of Antigravity to steer users toward Gemini. It didn’t. Google’s own frontier model, Gemini 4 Argon, launched just this week with a million-token window, so the company clearly isn’t short on models to promote. Adding Claude 5.5 anyway reads as a bet that developers stay for the workflow, not the logo on the model.

Who benefits most

Small teams and indie developers gain the most here. The flat subscription turns unpredictable per-token API spend into a known monthly cost, which is the difference between budgeting for AI coding help and hoping for the best. Startups that already run on Google Cloud can now route their agent-assisted development through infrastructure they already pay for, with the billing line item sitting next to their cloud bill instead of in a separate tab.

Enterprise teams get something subtler: a sanctioned place to compare models without a procurement process for each one. When Anthropic, Google and OpenAI all live behind one Google invoice, the security review covers the platform once instead of three times. That’s the kind of boring administrative win that actually decides which tools get adopted inside large companies.

The November 2 deadline is the actionable part

If you are reading this as someone who ships code with Antigravity, the thing to do this week is check which models your agents are pinned to. Anything on Claude 4.6 or GPT-OSS-120B needs a plan before November 2, and the 5.5 family is different enough in speed and cost that your prompts may behave differently under it. Run the comparison while both generations are still live in the dropdown. That’s the whole point of having them there.

Google turned its IDE into a model showroom without most people noticing. The dropdown is the feature. Use it like one.

Prince Mario-Max Schaumburg-Lippe: AI Giants Plan Joint Frontier AI Safety Standards Body

The world’s biggest AI labs are talking about doing the thing they’ve talked about for years: setting rules for themselves, together. According to The Information’s reporting published September 24, 2026, Google, OpenAI, and Anthropic are in discussions to create a joint body tentatively called the Standards Authority for Frontier AI (or SAFA), an industry-led organization that would set and enforce safety standards for the most powerful AI systems.

If it happens, it would be the most significant self-governance experiment in the history of the tech industry. And it would arrive at a moment when government-led AI regulation has mostly stalled.

What it would actually do

This wouldn’t be a talking shop, at least on paper. The functions under discussion:

  • Pre-deployment testing standards. Common requirements for evaluating frontier models before release, so “we tested it thoroughly” means the same thing at every lab.
  • Incident reporting. A shared framework for disclosing when AI systems malfunction or cause harm, something like how aviation or cybersecurity incidents get reported.
  • Auditor qualifications. Standards for who gets to audit AI systems and what counts as a rigorous audit, in a market that today ranges from serious to theatrical.

An OpenAI spokesperson has confirmed active talks with Google and Anthropic about coordinated safety frameworks. Whether SAFA should also run testing itself. The U.S. Center for AI Standards and Innovation (CAISI), which handles that job today, is widely seen as under-resourced for frontier systems. That’s still undecided.

These are precisely the gaps critics of AI self-regulation have pointed at for years. Voluntary commitments from individual labs are hard to compare, harder to verify, and easy to quietly abandon. A shared body with real definitions could change that, but only if the labs give it teeth.

The FINRA idea

The most intriguing detail is the institutional model. The body would reportedly be modeled on FINRA, Wall Street’s self-regulatory organization, an idea that traces to a proposal published July 14, 2026 by Sir Demis Hassabis, the head of Google DeepMind.

FINRA is not a government agency. It’s a private, industry-funded body with genuine enforcement power over broker-dealers, including the ability to fine firms and bar individuals. It works because participation is effectively mandatory for doing business in US securities markets, and because its rules have real consequences.

Translating that to AI raises obvious problems. FINRA’s authority ultimately rests on a statutory foundation: Congress built the framework that gives it power, and the SEC oversees it. An AI standards body with no government backstop would rely entirely on voluntary participation and reputational pressure. A draft White House executive order that would have brought federal supervision reportedly stalled, after the administration told the labs to find industry consensus first. Would OpenAI or Anthropic actually submit to binding judgments from a body their competitors co-founded? The history of tech self-regulation (social media moderation, privacy) says skepticism is the sane default.

Why the timing isn’t accidental

Federal AI safety efforts in the United States have stalled, leaving the most powerful technology of the century governed largely by the voluntary commitments of the companies building it. The labs seem to have concluded that waiting for legislation is no longer a strategy, and that shaping the rules themselves beats having rules imposed on them later. A credible industry standards body could also preempt heavier-handed government regulation, and regulators in the EU and elsewhere will be watching to see whether the body has substance or is mostly a shield against legislation.

Then there’s the pace of it all. Anthropic just shipped two frontier models in a single week. Anthropic’s new Sonnet 5.5 landed six days after Opus 5.5. And on September 23, OpenAI’s Sam Altman and Anthropic’s Dario Amodei addressed the United Nations Security Council to say the industry needs stronger oversight. When the labs are shipping this fast and appealing to the UN, “wait for the government” stops being a plan anyone believes.

The guest list

The CEO shortlist reportedly includes Sriram Krishnan, the former venture capitalist who served as senior White House AI policy adviser in the Trump administration, and Arati Prabhakar, the former director of the White House Office of Science and Technology Policy. Condoleezza Rice and venture capitalist David Friedberg have reportedly been approached for senior leadership roles.

Notice who these people are. Not AI researchers. People who understand Washington, institutions, and power. Krishnan and Prabhakar bring deep policy credibility; Rice brings geopolitical weight. The message is clear: this body wants to operate at the level of governments, not as a technical working group.

A launch is reportedly possible in late 2026 or early 2027. That’s an aggressive timeline that suggests the conversations are further along than a trial balloon. SAFA would succeed the Frontier Model Forum the same companies created in 2023.

Why it might work. Why it might not.

Start with the strong version. The three labs driving this represent the overwhelming majority of frontier AI capability. If they genuinely align on testing standards and incident reporting, that becomes the de facto global standard whatever anyone else does.

Now the weak version. Self-regulation serves the interests of the regulated. Standards written by the three biggest labs could easily become a moat: compliance costs that incumbents absorb without blinking but that crush open-source projects and smaller competitors. And without government enforcement, the ultimate sanction for violating the standards is disapproval. The history of tech self-regulation is littered with impressive-sounding bodies that produced impressive-sounding reports and changed very little.

There’s also a structural question nobody can dodge: who watches the standards body? If it’s funded by the labs, governed with lab input, and enforcing standards the labs wrote, its independence is inherently limited.

Nvidia’s Jensen Huang and Meta’s Mark Zuckerberg have publicly opposed the approach. And the research world is bigger than three labs. Serious, peer-reviewed advances are coming from unexpected places now: DOCOMO’s cold-start breakthrough is a telecom, not a frontier lab. Standards written only with the giants in the room will miss that.

Who else should pay attention

Startups and open-source developers: watch the auditor-qualification and testing standards closely. If these become industry norms, or get referenced in future regulation or procurement requirements, compliance costs could decide who can afford to build frontier-scale models. Don’t wait to be regulated by people you never met.

Enterprise buyers: a credible body would eventually let you compare vendors’ safety claims apples to apples. Start asking your vendors now how they test models pre-deployment and handle incident disclosure.

Policymakers: if governments want a seat at the table, the window is now, before the institution’s norms harden. Dismissing it as pure theater would be a mistake.

The Bottom Line

A joint Standards Authority for Frontier AI could be the moment the AI industry grew up institutionally, or an elaborate exercise in regulatory preemption. Which one it becomes depends on enforcement powers, funding independence, transparency, and whether anyone beyond the big three gets a real voice.

Prince Mario-Max Schaumburg-Lippe: Claude Sonnet 5.5: Anthropic’s New AI Workhorse Arrives

Anthropic’s shipping cadence is getting hard to keep up with. On Monday, September 28, the lab released Claude Sonnet 5.5, the second model in its Claude 5.5 family, arriving six days after the flagship Opus 5.5 launched on September 22. Opus is the showpiece. Sonnet is the engine room: the model most developers and businesses will actually run, day after day, at serious volume.

The timing is hard to ignore. Anthropic is releasing models at a clip its own CEO says the industry can’t sustain, and Reuters reports the company is preparing a Nasdaq IPO that could begin marketing as early as mid-October. Sonnet 5.5 sits at the center of all three stories.

The price didn’t move. The math did.

Sonnet 5.5 costs $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads, exactly the same as Sonnet 5. In a market where every generation usually arrives with a pricing tweak, standing pat is itself a statement.

But the sticker price isn’t the story. Efficiency is. Anthropic says the model needs far fewer tokens to complete the same work, costs up to 30% less for most work, and generates output more than 30% faster than its predecessor. At the scale these models run, where a single customer might push millions of API calls a month, that token efficiency compounds fast.

This is how frontier AI economics actually work now. The price per token matters less than the tokens required per unit of useful output. Anthropic is betting its customers can do that arithmetic. They’re probably right.

It can code. Really code.

The benchmark numbers deserve attention because they’re unusually decisive. On Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5.5 scored 70.6%: against 10.3% for Sonnet 5, and ahead of the flagship Opus 5.5’s 66.4% at its highest effort setting. On CursorBench and FrontierCode it similarly leapfrogged Sonnet 5. On the latter, scoring ten points above its predecessor at roughly one-fifteenth the task cost.

In Anthropic’s words, it’s “a faster, lower-cost complement to Claude Opus 5.5”: strongest at well-scoped everyday tasks: fixing bugs, creating polished documents, slides, and spreadsheets. It’s also the first Sonnet model to launch with frontier-grade cybersecurity safeguards and fallbacks comparable to the company’s most capable models, while its biology safeguards remain unchanged from Sonnet 5. That matters for enterprise procurement teams, who read safety posture as closely as they read benchmarks.

Three models, three jobs

The 5.5 lineup is now a clean ladder:

  • Opus 5.5 (September 22): the flagship, $4 input / $20 output per million tokens, built for the hardest reasoning, coding, and agentic work.
  • Sonnet 5.5 (September 28): the balanced workhorse at $2 / $10, faster and cheaper per task, good enough for the everyday heavy lifting.
  • Haiku 5.5: the lightweight speedster, due in the coming weeks, aimed at high-volume, cost-sensitive applications.

The positioning is unusually honest. Use Opus where quality is everything, Sonnet for the bulk of real workloads, Haiku where latency or cost dominates. It mirrors how cloud providers sell compute, which is no accident: it lets enterprise procurement teams slot models into tiers they already understand.

And it’s available everywhere on day one: the Claude Developer Platform (model ID claude-sonnet-5-5), AWS, Google Cloud, and Microsoft Azure. Existing cloud customers can adopt it without changing a thing. That ubiquity is a quiet weapon. It removes friction at the exact moment a team is deciding which model to standardize on.

Enterprise is the whole game

Here’s the number that explains Anthropic’s entire strategy: enterprise customers account for roughly 80% of the company’s business. The roster includes Salesforce, Databricks, Goldman Sachs, and Novo Nordisk, organizations that don’t experiment with AI so much as industrialize it.

Everything about Sonnet 5.5 reads like a product built for CIOs, not hobbyists. Token efficiency over benchmark bragging. Flat pricing. Day-one availability on every major cloud. Anthropic isn’t chasing the consumer chatbot crown; it’s building the model layer for corporate AI infrastructure, and Sonnet is the volume product. Even Meta’s enterprise push shows the rest of the industry has read the same memo.

Consumer AI is a brutal, low-margin attention business, and Anthropic lacks the distribution advantages of the giants. Enterprise rewards reliability, a safety reputation, and deep integration work, the things a research-first lab is actually good at.

The awkward essay

There is an irony here, and it deserves a straight look. On September 12, CEO Dario Amodei published an essay titled “We Must Pace the Frontier,” arguing the industry should slow the pace at which it improves AI capabilities. Sixteen days later, his company had shipped two frontier models in a single week.

Critics will call it hypocrisy. The fairer reading is that Amodei is describing a collective-action problem: no single lab can slow down alone without losing to competitors, so the fix has to be industry-wide coordination rather than individual restraint. Anthropic also says Sonnet 5.5 doesn’t advance the frontier of its models’ capabilities. This one is about efficiency, not a capability jump. And it helps that Anthropic is reportedly involved in the proposed joint safety standards body, exactly the kind of collective mechanism his argument would require.

Still, whether the “pace the frontier” rhetoric survives the quarterly pressure of a public listing is the thing to watch.

The IPO clock

Reuters reports Anthropic has picked Nasdaq for a potential IPO, with investor marketing possibly beginning in mid-October. Nvidia is reportedly in talks to invest as much as $10 billion as an anchor investor, at a discussed valuation in the region of $2 trillion. Read in that light, the 5.5 releases look like choreography: arrive at the roadshow with a fresh, complete lineup and a clean enterprise growth story.

It would be a landmark listing, arguably the first true frontier lab to go public, and it would put the company’s safety commitments under the fluorescent lights of public markets. Investors will want growth. The charter promises restraint. Sonnet 5.5 is the product that lets Anthropic claim both: growth through efficiency and adoption, not through ever-riskier capability jumps.

What to actually do with this

If you build on Claude: test Sonnet 5.5 against your current Sonnet 5 workloads before touching anything. The savings should show up in your bills within weeks, but verify quality on your own edge cases first.

If you’re picking a provider: map the Opus/Sonnet/Haiku ladder against your real workload mix. Most organizations overbuy capability; Sonnet 5.5’s efficiency gains might make previously-too-expensive workflows suddenly affordable. Worth an audit.

If you watch the industry: track the IPO. A public Anthropic will face quarterly pressure to grow API revenue, and enterprise adoption of efficient models is the healthiest way to do it.

The Bottom Line

Sonnet 5.5 isn’t a revolution. It’s something more useful: a better deal. Same price, fewer tokens per task, output more than 30% faster, available everywhere on day one, aimed at the enterprise customers behind 80% of Anthropic’s business. In a year of dramatic AI announcements, the releases that quietly make AI cheaper to run at scale will matter most. With an IPO reportedly weeks away, this one arrived right on schedule.

Prince Mario-Max Schaumburg-Lippe: GPT-6 vs Claude Opus 5.5: AI Model Price War Guide

On September 22, 2026, the AI industry witnessed something unprecedented: two frontier labs launched flagship models about 90 minutes apart, both slashing prices dramatically. Anthropic released Claude Opus 5.5, and OpenAI answered with GPT-6 Sol and GPT-6 Luna. API costs for top-tier AI just fell by roughly half — overnight.

If you pay for AI by the token, this is the best news you’ve had all year. Here’s what changed and how to take advantage.

The new lineup

Anthropic: Claude Opus 5.5

Launched September 22, Opus 5.5 is Anthropic’s new flagship, optimized for agentic coding and knowledge work. The headline numbers:

  • Pricing: $4 per million input tokens / $20 per million output tokens
  • Cost reduction: 40% cheaper than its predecessor while matching previous top-model performance
  • Positioning: the premium option for complex coding and long-horizon agent work

OpenAI: GPT-6 Sol and GPT-6 Luna

OpenAI split its release into two tiers — a clear segmentation play:

  • GPT-6 Sol: $2 per million input / $10 per million output — the workhorse, roughly half the cost of GPT-5.6-class models
  • GPT-6 Luna: $0.10 per million input / $0.50 per million output — the efficiency tier, aimed at high-volume production agents

The Sol/Luna split is strategically clever. Instead of one model trying to be everything, OpenAI is letting customers self-select: pay for quality where it matters, pay pennies where it doesn’t.

Head-to-head: benchmarks

For data science and engineering workloads, the numbers favor Anthropic at the top end:

  • Terminal-Bench 4.0: Opus 5.5 scores 66.4% vs. GPT-6 Astra’s 57.9%
  • Frontier Code v1.1: Opus 5.5 scores 54.4% vs. Astra’s 53.3%
  • Cost per task: Opus 5.5 runs roughly 60% cheaper than Astra for equivalent coding work

But benchmarks only tell part of the story. GPT-6 Astra still leads in frontier math, science, and abstract reasoning — the kind of work where raw capability matters more than cost per token. And Luna’s pricing is so aggressive ($0.10/$0.50) that for high-volume, simpler tasks, nothing else is close.

Which model for which workload

Here’s a practical decision framework:

Choose Claude Opus 5.5 when:

  • You’re doing agentic coding — multi-step refactors, test-driven development, codebase-wide changes
  • You need long-context analysis of documents or data
  • You’re running knowledge-work agents where quality compounds (research, analysis, writing)
  • Your bottleneck is capability, not budget

Choose GPT-6 Sol when:

  • You need strong general performance at moderate cost
  • You’re running production agents at meaningful volume
  • You want the best balance of quality and price for mixed workloads

Choose GPT-6 Luna when:

  • You’re processing high volumes of simpler tasks — classification, extraction, summarization
  • You’re building features where per-unit economics make or break the product
  • You need “good enough” intelligence at massive scale

The smartest move for most teams: route dynamically. Use Luna for the 80% of tasks that are routine, Sol for the 15% that need real judgment, and Opus 5.5 for the 5% where quality is everything. The price gaps are now large enough that intelligent routing can cut your AI bill by 70% or more without visible quality loss.

Why prices are falling

This isn’t charity — it’s competition, and it’s coming from three directions:

  1. Open-weight pressure. Chinese models now account for over half of usage on OpenRouter, a major developer platform. Alibaba’s Qwen-Audio-3.1 launched with up to 95% API price reductions. When capable open models are nearly free, closed labs have to justify every dollar.
  2. Efficiency gains. Both releases emphasize doing more with less compute. Architectural improvements mean the same hardware now serves more tokens — and labs are passing some of those savings on to win market share.
  3. The agent land grab. Every lab wants developers building agents on their platform, because agents create sticky, high-volume API usage. Cheap tokens are customer acquisition cost. Meta’s enterprise Muse platform, Microsoft’s Copilot overhaul, and OpenAI’s rumored persistent assistant all point to the same bet: win the developer, win the decade.

What this means for your AI budget

Renegotiate now. If you’re on committed-use contracts priced against older models, the market just moved. The 40-50% reductions are public and immediate — use them as leverage.

Revisit “too expensive” projects. AI features that didn’t pencil out six months ago might work today. That support agent, document pipeline, or code assistant that was 2x over budget? Run the numbers again with Luna or Sol pricing.

Watch for the next shoe. Google’s Gemini 3.8 is expanding across agentic applications, and the synchronized timing of these releases suggests the labs are watching each other closely. Another round of cuts before year-end wouldn’t surprise anyone.

Don’t chase price alone. The cheapest model that does the job is the right model — but “does the job” needs testing, not assumptions. Run your actual workloads against two or three options before committing. A model that’s 90% cheaper but produces 20% more errors can cost more in the end.

Bottom line: The frontier AI price war just made powerful models dramatically cheaper. For builders, this is a golden window — capabilities that were premium-priced last month are now commodity-priced. The winners won’t be the teams with the biggest AI budgets, but the teams that route intelligently across a suddenly diverse model landscape.