Overview
The dominant theme in today's AI news is a widening gap between capability and control. Anthropic's newest Claude model, Opus 4.6, apparently crumbled under even amateurish jailbreak attempts designed to generate sexually explicit content — an awkward moment for a lab whose entire identity rests on safety-first development. Security researchers also demonstrated successful data-exfiltration attacks against both Grok and Microsoft Copilot, while Meta found itself in the crosshairs over ads promoting a non-consensual deepfake nudity app. For an industry desperate to be trusted by enterprises and regulators alike, it was not a great day.
But there's a constructive side to the story. Nvidia research out today suggests the "harness" around an AI model — the fine-tuning and orchestration that guides it — matters more than raw model capability when it comes to keeping agents on task. That is a genuinely important reframe for anyone buying AI systems. Meanwhile, the infrastructure arms race shows no signs of cooling: Nvidia is pouring more money into data centers through a new Cloverleaf partnership, Starcloud raised fresh capital for orbital data centers, and WIRED spotlights how a city in Inner Mongolia has become an unlikely hub for China's AI boom thanks to cheap energy, abundant land, and proximity to Beijing.
On the model front, Grok 4.6 topped CursorBench 3.2 at a fraction of the cost of its rivals — a sign that efficiency is becoming a competitive weapon. For anyone tracking the stack from chips to chatbots, GetAI Business remains a solid home base for discovering the tools reshaping this fast-moving landscape.
Today's Big News
Anthropic's Opus 4.6 Is Shockingly Easy to Jailbreak
TechCrunch found that bypassing Claude's restrictions on explicit content took surprisingly little effort, despite Anthropic's famously strict usage policies. For a company that has built its brand on safety leadership, this is a credibility problem — and a reminder that refusal training is still an arms race, not a firewall.
Nvidia Proves the Harness Beats the Model — and Invests Accordingly
Nvidia's research shows that with the right fine-tuning and orchestration, AI agents perform well — and stay out of trouble — even when the underlying model is mediocre. The company also announced a partnership with data center developer Cloverleaf, deepening its investment in the physical infrastructure that powers AI. It's a two-pronged bet on both the software layer and the silicon beneath it.
Grok and Microsoft Copilot Both Hit by Data-Theft Attacks
Researchers demonstrated a novel "cryptographic context injection" attack against Grok, exfiltrating user data when malicious instructions were encrypted — the latest twist in breaking LLM guardrails. Separately, a secret parameter in Microsoft Copilot allowed hackers to steal passwords after a target clicked a link. Two of the most widely deployed AI assistants, two security holes, one very busy day for red teams.
Meta's Trust Deficit Keeps Growing
Meta ran ads for an app promising to "nudify" female politicians, with at least one ad featuring a deepfake closely resembling a US politician. At the same time, the explosive growth of Meta AI glasses is fueling privacy backlash — and the latest detection app, Zuckoff, is itself imperfect. Meta is simultaneously the consumer AI leader and the platform most out of step with public trust.
Starcloud Raises $250M to Take Data Centers to Orbit
Starcloud's new round targets orbital data centers as terrestrial power and land constraints mount — but with launch options drying up, a serious fight over access to space is brewing. Back on Earth, the buildout is equally frenetic: Nvidia's Cloverleaf tie-up and Inner Mongolia's emergence as an AI hub prove that gravity hasn't slowed the infrastructure boom one bit.