Prince Mario-Max Schaumburg-Lippe: Prime Intellect Launches AI Inference for Open Models

Who Controls the Pipes Wins

Announced October 3, 2026 — the freshest launch in this week's AI news cycle — Prime Intellect publicly launched Prime Inference, a serving platform for frontier open-source models. It is the serving layer of the company's open training stack, sitting alongside post-training tools like prime-rl, verifiers, and sandboxes.

The pitch is straightforward: two modes of serving. Serverless endpoints for variable demand, reserved capacity for sustained workloads. Both run on Prime's own GPUs across multiple datacenters — NVIDIA Blackwell hardware today, with Vera Rubin listed as coming soon.

But the launch metrics are what make this announcement land. Before going public, Prime Intellect says it processed nearly a trillion tokens per day internally — across RL rollouts, synthetic data generation, evaluations, and coding agents — with a near-zero tool-call error rate and 100% uptime since launch. Its GLM-5.3 endpoint is reported among the fastest on OpenRouter. This is not a pitch deck; it's a load-tested system being opened to the public.

The New Moat Isn't the Model

The 2026 AI story is quietly shifting. For years the industry fixated on who builds the best model. But there's a growing realization that the real leverage sits one layer down: who controls the pipes between models and users.

Prime Intellect's thesis is that teams should be able to train, evaluate, and serve their own models on their own stack — "own their intelligence" rather than depend on frontier providers. The technical details show what that takes in practice: an OpenAI-compatible API (point any OpenAI SDK at https://api.pinference.ai/api/v1), automatic failover across datacenters, cache-aware routing that uses host DRAM as a secondary KV tier for large prompts, and targets around 100 tokens per second. Wisevoter's launch coverage notes the platform serves GLM-5.3 on GB200 NVL72 hardware with OpenAI-compatible SDKs.

That last part matters. Inference optimization — cache-aware routing, dedicated Blackwell capacity — is where AI economics are being won right now, not just in model benchmarks. Every percentage point of serving efficiency is a percentage point of margin, and at trillion-token scale, those points are worth real money.

The Closed Loop

The deeper idea here is the closed loop. Prime Intellect isn't just selling inference; it's building an integrated compute, training, inference, and sandbox stack where production traces feed back into training. Serve the model, watch how it's used, fold the data back into the next training run. That's the flywheel that compounds — and it's the same pattern driving the industry's rush to squeeze more compute from existing power and the massive data center buildouts feeding the open-model ecosystem.

There's also a community angle that shouldn't be overlooked. Prime Intellect is backed by Founders Fund, Radical, NVIDIA, Intel, and AI researchers including Andrej Karpathy and John Schulman. In the open-source AI world, credibility is currency, and those names spend. The open-model community has been waiting for production-grade serving infrastructure that isn't controlled by the frontier labs — Prime Inference is a serious answer to that wait.

Why Tool-Call Reliability Is the Real Spec

Buried in the launch metrics is the number that matters most for where AI is heading: a near-zero tool-call error rate. That spec isn't about chatbots. It's about agents.

Coding agents, RL rollouts, synthetic data generation — the workloads Prime Intellect ran before launch — all depend on models that can call tools reliably, thousands of times in a row, without a malformed call breaking the chain. One failed tool call in a hundred might be fine for a demo. At a trillion tokens a day, it's a catastrophe. The fact that Prime load-tested on exactly these workloads suggests the platform was built for the agentic era from day one, not retrofitted for it.

The target of around 100 tokens per second is the other number worth pausing on. For interactive use, that puts it in the territory where an open model served on Prime's stack feels roughly as responsive as a closed API. And because the API is OpenAI-compatible, switching costs are low: point an existing SDK at a new URL and you're running. That's how you win developers — not with a migration guide, but with a one-line config change.

What to Watch

The open question is scale economics. Running dedicated Blackwell capacity across multiple datacenters is expensive, and the inference market is brutally competitive. Prime's advantage is vertical integration: the same stack that serves models also trains them, which means optimizations can flow in both directions. If the closed loop works — production traces genuinely improving the next model generation — that integration becomes a moat that pure serving providers can't easily copy.

For teams building on open models, though, the practical takeaway is simpler. A fast, reliable, OpenAI-compatible serving layer for frontier open models lowers the cost of independence. You don't have to choose between the convenience of a closed API and the freedom of an open model anymore. The pipes are getting built, and they're getting cheaper.

The Takeaway

The 2026 AI battle is moving from models to infrastructure — and inference is the new front line. Prime Intellect's near-trillion-token launch track record gives Prime Inference instant credibility, and its closed-loop stack points at where the industry's economics are heading. Own the pipes, and the models follow.

Prince Mario-Max Schaumburg-Lippe: Karman, Aolani Unlock 50% More AI Compute From Same Power

While everyone else hunts for more megawatts, two companies just announced a way to get up to 50% more AI compute out of the power that’s already there. On October 1, Singapore-founded neocloud Aolani announced a partnership with Karman (formerly Utilidata) to deploy advanced power orchestration across Aolani’s AI infrastructure — and the headline number is the strongest data-center efficiency claim of the week: 50% more compute capacity from the same provisioned power.

No new substations. No grid-connection queue. No waiting. Just smarter use of the electrons already flowing.

How it works

Karman’s platform pairs high-resolution power metrology with local processing and AI, running on a custom NVIDIA Jetson Orin Nano at the rack level. In plain terms: it measures exactly how much power each rack of GPUs is actually drawing, moment to moment, and dynamically reallocates the available capacity in real time. Data centers provision power for worst-case peaks that rarely arrive simultaneously; the gap between provisioned and used is “stranded power,” and it’s enormous.

Karman says the platform increases tokens per watt by 50% by unlocking that stranded capacity. This isn’t a lab demo. At its first commercial deployment in North America, the system already unlocked 33% more compute capacity from existing power infrastructure — which the company says could translate into roughly $20 million in additional revenue per megawatt. The joint proof-of-concept with Aolani will run on the NVIDIA Blackwell platform.

Why “do more with what you have” is the story of 2026

Aolani CEO Nicholas Chia put it simply: the goal is to “deliver the compute capacity quickly without waiting on the grid.” That sentence is the whole ballgame. The AI industry’s binding constraint has shifted from silicon to substations. A new hyperscale data center can take two to four years to get connected to the grid in some markets. Efficiency software that adds 50% capacity overnight is, functionally, a time machine.

It pairs neatly with the other side of this week’s news: neoclouds pledging GPUs as collateral to finance new builds and Japan co-locating data centers with power plants. The industry is attacking the power problem from both ends — more supply and better utilization. This announcement is the purest expression of the second approach.

The economics flip

Here’s the part the finance people will circle. If a software layer can lift a data center’s effective capacity by a third to a half, power efficiency stops being an operations concern and becomes a revenue line. Twenty million dollars per megawatt of unlocked capacity is a number that reorders priorities fast. Suddenly the efficiency team is the growth team.

And there’s a climate angle that deserves a mention. Every percentage point of utilization gained from existing infrastructure is capacity that doesn’t need a new power contract — or a new power plant. In a year when data-center electricity demand is straining grids and climate commitments simultaneously, doing more with the same electrons is the rare win that pleases the CFO and the sustainability officer at once.

Power is the new compute. The companies that treat it that way — measuring it, managing it, squeezing it — are building the real infrastructure of the AI era. Karman and Aolani just showed everyone the math.