Skip to main content

Tech & AI News — 28 June 2026

· 13 min read

OpenAI unveiled its first custom AI chip, Anthropic accused Alibaba of running the largest-ever AI theft campaign using 25,000 fake accounts, and Google launched a reasoning model that outscored every public competitor on the hardest science benchmarks — all in the same week that SpaceX's Tennessee data centres quietly became the AI world's busiest compute exchange. If you want to talk about where AI infrastructure power is pooling, and who is trying to steal it, this is the week that wrote the case study.

This week in one line

OpenAI entered the chip race, Alibaba allegedly ran 25,000 fake accounts to steal Claude's brain, and Google's new reasoning engine topped every science benchmark on the board.

The big stories

OpenAI and Broadcom unveil Jalapeño — OpenAI's first custom AI chip

On 24 June, OpenAI and Broadcom jointly announced Jalapeño, OpenAI's first custom-designed AI accelerator. The chip is purpose-built for inference — the process of serving a trained model's responses to users in real time — rather than for training new models from scratch. That distinction matters: inference is the part of the AI stack that scales with every ChatGPT query, every API call, every Codex completion.

Broadcom CEO Hock Tan said Jalapeño delivers roughly 50% cost savings compared with standard AI GPUs. OpenAI President Greg Brockman told CNBC the chip was designed end-to-end in nine months, with AI models helping write parts of the design — possibly the fastest ASIC development cycle in the history of high-performance semiconductors. Engineering samples were physically delivered to OpenAI leadership on the announcement day. Broadcom handles silicon fabrication and networking; OpenAI designed the accelerator architecture.

The first deployment is targeted for late 2026. Jalapeño is framed as the first generation in a multi-generation compute platform — OpenAI's long-game bet that owning silicon lets it escape perpetual dependence on Nvidia's pricing and production allocation.

Why it matters globally: every frontier AI lab — Anthropic, Google, Meta, OpenAI — currently competes for the same pool of Nvidia GPUs at prices Nvidia sets. A lab that can manufacture its own inference chips at half the cost changes its unit economics and its negotiating position. If Jalapeño hits its cost targets, OpenAI's per-query margin improves — and that eventually feeds into pricing, capacity, and the ability to deploy more capable models at scale.

South Africa / Africa relevance: African developers and businesses pay for AI API calls in USD, with Nvidia GPU costs baked in. If the next generation of inference hardware is 50% cheaper to run, that cost reduction should eventually flow through to API pricing — lowering the ZAR barrier for SA developers building on Claude, GPT, or any model run on Jalapeño-class silicon. The harder question is whether African compute infrastructure will ever get access to custom chips, or whether local data centres will remain dependent on commodity GPUs that the big labs no longer prioritise.


Anthropic accuses Alibaba's Qwen lab of the largest AI theft campaign on record

In a letter to the US Senate Banking Committee dated 10 June 2026, Anthropic accused operators linked to Alibaba and its Qwen AI lab of conducting "the largest known distillation attack on Anthropic to date." Tom's Hardware confirmed the key numbers: approximately 25,000 fraudulent accounts made 28.8 million exchanges with Claude between 22 April and 5 June 2026.

A distillation attack works like this: you feed thousands of carefully designed prompts to a frontier model, collect its detailed outputs, and use that input-output dataset to train your own smaller model to behave like the original. You are, in effect, extracting the capabilities of a model you don't own. It is legally contested territory — OpenAI raised the same concern about DeepSeek earlier this year — but Anthropic's letter to the Senate names Alibaba explicitly and asks for legislative action.

To get past Anthropic's geographic restrictions (Chinese entities cannot access Claude under US export rules), the accounts used commercial proxy services to mask their origin. The scale is notable: the Alibaba campaign was 1.7 times larger than the combined February 2026 campaigns attributed to DeepSeek, Moonshot AI, and MiniMax, which themselves used 24,000 fake accounts and 16 million exchanges. Alibaba has denied training on proprietary model outputs.

Congress is now preparing legislation to sanction Chinese AI rivals in response to Anthropic's filing.

Why it matters globally: this story marks an escalation in the AI IP war from trade dispute territory to something closer to active intelligence operations. The scale — 25,000 fake accounts running for six weeks — suggests an organised, well-resourced effort, not a side project. It also reveals how porous the API access layer is as a security boundary: geographic restrictions are only as strong as proxy detection.

South Africa / Africa relevance: any SA developer or company that builds a product on top of AI APIs should understand what distillation attacks mean for the ecosystem. If large actors can effectively copy frontier models cheaply, the commercial moat of those APIs narrows — which could reduce the labs' incentive to maintain affordable API pricing, or push them toward more restrictive access controls that affect legitimate SA users alongside bad actors.


Google Gemini 2.5 Pro with Deep Think tops every public science benchmark

On 22 June, Google launched Gemini 2.5 Pro with Deep Think reasoning mode for Google AI Ultra subscribers. The model scored 82.4% on GPQA Diamond — a benchmark of PhD-level questions in chemistry, physics, and biology — surpassing Anthropic's Claude Fable 5 (79.1%, which has been offline since the government-ordered suspension) and OpenAI's GPT-5.5 (76.3%). On MMLU-Pro, a broader professional-knowledge benchmark, Gemini 2.5 Pro Deep Think scored 89.8% — the highest ever recorded by a publicly available model.

Deep Think is Gemini 2.5 Pro's extended reasoning mode. Rather than returning an answer after one pass through the model, Deep Think explores multiple reasoning paths in parallel across more iterations before settling on a response — a compute-intensive approach that trades response speed for accuracy on hard questions. It sits inside the same Gemini 2.5 Pro model; you enable it by selecting the "Deep Think" option in the interface or API.

Deep Think is currently available to Google AI Ultra subscribers, with API access for developers described as "arriving soon." Gemini 3.5 Pro, a separate and more powerful upcoming model, has been delayed to July 2026 while Google continues internal testing.

Why it matters globally: the benchmark result is notable because Fable 5's absence from the leaderboard (due to the US export control suspension covered in last week's edition) leaves an opening, and Google has stepped into it with hard numbers. The race at the frontier is no longer about releasing a new model every few months — it is about which model you can actually access today. Gemini 2.5 Pro Deep Think is live. Fable 5 is not.

South Africa / Africa relevance: Google AI Pro and Ultra plans are accessible from South Africa. SA developers and researchers who need serious scientific or legal reasoning — complex contract analysis, research synthesis, technical documentation — have a live, publicly accessible option that currently tops the leaderboard. Worth experimenting with before the Gemini 3.5 Pro release shifts the comparison again.


SpaceX's Colossus data centres become the AI world's compute marketplace

On 22 June, SpaceX signed a $150 million per month compute deal with Reflection AI, an open-weight AI startup backed by Nvidia. Reflection gets access to Nvidia GB300 GPUs at SpaceX's Colossus 2 facility in Memphis, Tennessee, starting 1 July 2026. The contract runs through 2029 and is worth up to $6.3 billion. Either party can exit after the first three months with 90 days' notice.

The deal means Colossus is now leased to at least four external AI organisations: Anthropic (paying $1.25 billion per month), Google, Cursor, and now Reflection AI. SpaceX's xAI has quietly transformed the facility — originally built to train Grok — into a neutral compute provider, renting capacity to its competitors. Reflection is building open-weight frontier models it plans to release publicly, positioning itself as an open-source alternative to Anthropic and OpenAI.

Bloomberg confirmed the deal terms and the Colossus 2 location. TechCrunch noted Reflection framed the deal as proof that open-weight AI can compete with closed labs when given equivalent compute.

Why it matters globally: SpaceX has accidentally created the most influential private AI compute exchange on the planet. When Anthropic, Google, Cursor, and Reflection are all renting compute from the same Elon Musk facility in Tennessee, the infrastructure layer of the AI industry is more concentrated — and more dependent on one man's data centres — than the labs publicly acknowledge.

South Africa / Africa relevance: Africa's AI future depends on access to large-scale compute. The Colossus deals show that even trillion-dollar companies lease compute rather than build their own. SA and African cloud providers face a growing gap: not just in GPU availability, but now in the kind of custom, high-density GB300 infrastructure that only a handful of facilities worldwide can run. The Servernah Cloud sovereign AI platform launched in Nairobi in March 2026 is the type of local answer the continent needs — but it is orders of magnitude smaller.


Broader tech highlights

  • ChatGPT's market share fell below 50% for the first time, dropping to 46.4% in late May as Google's Gemini climbed to 27.7% and Claude reached 10.3% of AI assistant queries. The figures, reported across multiple analytics outlets, reflect a market that is genuinely multi-polar after three years of near-monopoly.

  • Colorado's AI Act is now delayed to 2027. The law — which required deployers to run risk assessments and disclose algorithmic decision-making — was originally set to take effect on 30 June 2026. The governor signed an amendment in June shifting the new effective date to 1 January 2027 and significantly narrowing the obligations, abandoning the EU-style risk-management framework in favour of narrower transparency disclosures. Hunton Andrews Kurth has the legal analysis.

  • OpenAI acquires Ona (formerly Gitpod) for persistent Codex agents. Announced 11 June, the Kiel, Germany-based startup gives Codex the ability to run coding tasks for hours or days in sandboxed cloud environments after a developer closes their laptop. Codex now has 5 million weekly active users, up from 3 million in April.

South Africa & Africa angle

  • SA AI policy redraft continues. Minister Malatsi's seven-member expert panel (convened after the hallucinated-citations scandal covered in last week's edition) is working toward a replacement draft for public comment by the end of 2026. The comment period, when it opens, is a genuine opportunity for SA civil society and tech sector voices to shape the framework that governs how AI is deployed locally.

  • Alibaba theft story has SA implications. South African companies building products on Claude, GPT, or other APIs should note that the Anthropic-Alibaba case establishes "distillation attack" as a recognised IP threat category. If labs respond by tightening access controls globally — stricter geographic checks, rate limiting, mandatory KYC — SA developers without a formal API relationship could face friction. Registering directly with Anthropic and Google Cloud (rather than going through resellers) is the lower-risk path.

  • Africa's compute gap deepens. As Colossus deals stack up — Anthropic, Google, Cursor, Reflection all leasing from SpaceX in Tennessee — the geographic concentration of frontier AI compute moves further away from Africa. The continent currently holds roughly 0.6% of global data-centre capacity while representing 18% of the world's population, as the Al Jazeera investigation reported on 26 June. South Africa leads the continent with about 40% of regional data-centre space; Johannesburg has 15 active facilities with six more under construction. Local ownership remains the open question.

Skills & learning corner

  1. Try Gemini 2.5 Pro with Deep Think on a hard reasoning problem this week. Google AI Ultra gives access now. Compare it on a legal, scientific, or financial question where you care about step-by-step accuracy — the benchmark score translates to real-world tasks in those domains.

  2. Read Anthropic's platform docs on Claude Fable 5 — even while the model is suspended. The Fable 5 developer guide explains the new refusal-handling API, fallback mechanism, and billing rules you will need when access is restored. Getting this wired up now means your integration is ready on day one.

  3. Understand what a distillation attack is. If you build on AI APIs professionally, you should be able to explain this to a client or employer. The CNBC Alibaba story is a readable primer. The short version: you can train a weaker model to behave like a stronger one by harvesting its outputs at scale — which is why API providers increasingly care about usage anomalies.

  4. Watch Reflection AI's open-weight model strategy. With $6.3B in compute locked in and Nvidia backing, Reflection is positioning as the open-source frontier alternative to Anthropic and OpenAI. If they ship a competitive open-weight model, SA developers could run powerful AI locally without API costs or export-control risk. Follow their progress at reflectionai.com.

OpenAI + Broadcom Jalapeño chip

Anthropic vs Alibaba distillation attack

Gemini 2.5 Pro Deep Think

SpaceX Colossus compute deals

Colorado AI Act delay

OpenAI acquires Ona

South Africa & Africa


Researched and verified on 28 June 2026. Primary sources were confirmed via search metadata and direct URL access where available. Direct fetch access was partially limited by the remote environment's network proxy this edition — all cited outlets are recognised primary or tier-2 sources per the editorial standard.