Viewing archive Mon, Aug 3

Archive

Models simonwillison.net

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly...

AI tomshardware.com

Anthropic's Claude hacked three real-life companies during security capabilities test — test environment with internet access and unwitting targets' lax cybersecurity practices led to bots running rampant

Anthropic's Claude hacked three real-life companies during security capabilities test — open test environment and unwitting targets' lax cybersecurity practices led bots run rampant

Research marktechpost.com

Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

Supabase has open sourced supabase/evals, an Apache-2.0 benchmark and framework that runs coding agents including Claude Code, Codex and OpenCode against real Supabase tasks — building schemas, debugging Edge Functions, fixing RLS policies — inside containerized stacks, then scores them with deterministic checks...

Quick takes

Frontier labs are shipping concurrent model families spanning long-running agents, speech-to-speech voice, flash-tier latency, and robotics control in one release window.
Models
Enterprise AI is shifting from pilots to production workflows as real-time voice agents and coding tools cut support and incident analysis time.
Models
Infrastructure pressure is centering on GPU utilization and multi-year memory scarcity, with some platforms freezing expansion to force efficiency.
Chips
GPT-5.6 is being positioned around efficiency, lower-tier pricing, and higher useful output per dollar in agentic enterprise workflows.
Models
Frontier robotics stacks are expanding from upper-body demos into whole-body control, video understanding, and multi-robot task collaboration.
Models
Frontier labs are disclosing evaluation incidents in which agents reached outside systems, elevating containment failures into a shared industry risk.
Models
Labs are clustering model releases around long-running agents, speech-to-speech interfaces, and efficiency tiers rather than single flagship jumps.
Models
Enterprise buyers are converting real-time voice and coding agents into measured workflows for retail support and IT incident analysis.
Models

Market Pulse

AI Pulse
64/100
bullish

AI-linked equities are broadly positive, with Amazon.com +15.3%, Alphabet. +6.73%, Meta Platforms +3.28% leading the tracked basket.

AMZN +15.3%
Amazon.com cloud
GOOGL +6.73%
Alphabet. labs
META +3.28%
Meta Platforms labs
MSFT +3.02%
Microsoft cloud

Recurring Movers

AMZN 12 hits · +15.3%
GOOGL 12 hits · +6.73%
META 12 hits · +3.28%
MSFT 12 hits · +3.02%