Viewing archive Sun, Aug 2

Archive

Models simonwillison.net

Ten advances in mathematics and theoretical computer science

Ten advances in mathematics and theoretical computer science A few days ago it was Anthropic discovering cryptographic weaknesses with Claude using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly...

Research marktechpost.com

Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

Supabase has open sourced supabase/evals, an Apache-2.0 benchmark and framework that runs coding agents including Claude Code, Codex and OpenCode against real Supabase tasks — building schemas, debugging Edge Functions, fixing RLS policies — inside containerized stacks, then scores them with deterministic checks...

Models theheadandtale.com

Money & Machines: Postpaid is back. The golden era isn't; What Sarvam's Chaplot bet signals

In today's Money & Machines edition, we unpack Paytm's renewed Postpaid push and what has changed since its heyday; and on the AI front, we examine what Sarvam AI's appointment of frontier AI researcher Devendra Chaplot signals as the startup sets out to build a trillion-parameter foundation model and expand its...

Models lesswrong.com

SOTA alignment assessments don’t strongly update US against misalignment

Anthropic concluded in the April Mythos Preview alignment risk update that the model "does not possess any unknown propensities that would increase alignment risk." The report argues that if Mythos Preview were coherently misaligned [1] [2], it likely would have been detected by the assessment (following Anthropic,...

Quick takes

Robotics development is shifting toward integrated systems that combine video understanding, whole-body control, task orchestration, dexterity, simulation, and multi-robot collaboration.
Models
AI infrastructure pressure is moving attention from raw accelerator capacity toward utilization, CPU inference, quantization, and useful intelligence per dollar.
Chips
Agentic systems are creating a new security evaluation burden: testing must account for internet access, unauthorized actions, containment failures, and ordinary defensive controls.
Policy
Recent releases span voice, long-running agents, coding, lightweight variants, cybersecurity models, and lower-cost inference, broadening deployment options.
Models
Enterprise adoption is increasingly framed around operational workflows, with voice, chat, incident analysis, workforce enablement, and multilingual support as deployment contexts.
Models
Governance is expanding from safety and provenance practices toward legal responsibility for agent behavior, supply-chain designations, and delegated financial tasks.
Policy
Recent model releases are broadening AI capability across robotics, voice interaction, long-running agents, and specialized model variants, expanding deployment decisions.
Models
Robotics development is shifting toward embodied systems that combine video understanding, whole-body control, task orchestration, and multi-robot collaboration.
Models

Market Pulse

AI Pulse
64/100
bullish

AI-linked equities are broadly positive, with Amazon.com +15.3%, Alphabet. +6.73%, Meta Platforms +3.28% leading the tracked basket.

AMZN +15.3%
Amazon.com cloud
GOOGL +6.73%
Alphabet. labs
META +3.28%
Meta Platforms labs
MSFT +3.02%
Microsoft cloud

Recurring Movers

AMZN 12 hits · +15.3%
GOOGL 12 hits · +6.73%
META 12 hits · +3.28%
MSFT 12 hits · +3.02%