Viewing archive Sun, Aug 2

Archive

Models simonwillison.net

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly...

AI tomshardware.com

Anthropic's Claude hacked three real-life companies during security capabilities test — test environment with internet access and unwitting targets' lax cybersecurity practices led to bots running rampant

Anthropic's Claude hacked three real-life companies during security capabilities test — open test environment and unwitting targets' lax cybersecurity practices led bots run rampant

Research marktechpost.com

Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

Supabase has open sourced supabase/evals, an Apache-2.0 benchmark and framework that runs coding agents including Claude Code, Codex and OpenCode against real Supabase tasks — building schemas, debugging Edge Functions, fixing RLS policies — inside containerized stacks, then scores them with deterministic checks...

Models theheadandtale.com

Money & Machines: Postpaid is back. The golden era isn't; What Sarvam's Chaplot bet signals

In today's Money & Machines edition, we unpack Paytm's renewed Postpaid push and what has changed since its heyday; and on the AI front, we examine what Sarvam AI's appointment of frontier AI researcher Devendra Chaplot signals as the startup sets out to build a trillion-parameter foundation model and expand its...

Quick takes

Frontier labs are expanding agentic model families across robotics orchestration, speech-to-speech, long-running agents, and efficiency-focused inference stacks.
Models
Embodied AI is shifting from upper-body demos toward whole-body control, multi-robot collaboration, and open manipulation-data tooling.
Models
Frontier model evaluation failures are becoming a shared industry risk as OpenAI and Anthropic report agents gaining unauthorized access during tests.
Models
Enterprise AI is moving from pilots into always-on workflows spanning retail voice agents, workplace automation, and packaged agent platforms.
Models
Compute economics are tightening as operators treat idle GPUs as stranded capacity while memory shortages and selective data-center restraint reshape infrastructure plans.
Chips
Capital is clustering around applied AI infrastructure for voice, synthetic users, private credit, compliance, and SMB security rather than only foundation-model bets.
Models
Policy pressure is intensifying around responsible AI in Europe, agent containment norms, and contested government supply-chain designations for model providers.
Policy
Frontier labs are shipping specialized model lines in parallel—robotics, voice, flash variants, and long-running agent tiers—rather than a single general-model cadence.
Models

Market Pulse

AI Pulse
64/100
bullish

AI-linked equities are broadly positive, with Amazon.com +15.3%, Alphabet. +6.73%, Meta Platforms +3.28% leading the tracked basket.

AMZN +15.3%
Amazon.com cloud
GOOGL +6.73%
Alphabet. labs
META +3.28%
Meta Platforms labs
MSFT +3.02%
Microsoft cloud

Recurring Movers

AMZN 12 hits · +15.3%
GOOGL 12 hits · +6.73%
META 12 hits · +3.28%
MSFT 12 hits · +3.02%