AiHub.

The ledger

240 ranked

AI agents being tested by OpenAI involved in cyber-attack on another service, say researchers

Two months before hacking Hugging Face, malicious packages authored by internal OpenAI agents were uploaded to RubyGems Agents being tested by OpenAI uploaded hundreds of malicious packages in a cyberattack on software service RubyGems in May, two ⁠months ​before they hacked open-source platform Hugging Face, the...

Big Tech  theguardian.com    signal 65  verified

Signals

All →

Dev tool ChatGPT

Greg Brockman cites 5M ChatGPT sites and new feature drops

OpenAI’s Greg Brockman (@gdb) posted that ChatGPT sites have reached 5M and that new feature drops are accompanying that milestone. Specific feature contents are not named in the supplied text.

observed

2 sources
  • @gdb — 5M ChatGPT sites, and new feature drops:
  • @thsottiaux — Massive upgrade to creating, sharing and hosting sites directly through ChatGPT.

Model release Smaug Flash

Bindu Reddy announces Smaug Flash open-weights fine-tune

@bindureddy announced Smaug Flash as an open-weights fine-tune aimed at personal agents, saying it is available on the RouteLLM API at $0.10M input and $0.40M output. Authority is the poster’s own claim; no other supplied posts corroborate it.

observed

1 source
  • @bindureddy — 🚨 Announcing Smaug Flash - Open Weights, Dirt Cheap And Tuned For Personal Agents Excited to announce Smaug Flash, our latest open-weights fine-tune. Available on the RouteLLM API - $0.10 M input

Dev tool Claude Code

Claude Code adds plugin eval for testing plugin value

@claudedevs announces Claude Code’s claude plugin eval: create test cases, run a plugin or skill against them, score the runs, and compare by rerunning without the plugin.

observed

1 source
  • @claudedevs — New in Claude Code: claude plugin eval See what value your plugin is adding, or if it needs more work. You can create test cases, run your plugin or skill against those test cases, score those runs

Hot OpenAI

GPT Images 2.5 tops Artificial Analysis Image Arena

@artificialanlys reports OpenAI’s GPT Images 2.5 taking the top two Image Arena spots, with higher quality in under half the time at the same price as GPT Image 2. This is a third-party benchmark result, not a primary release notice.

observed

1 source
  • @artificialanlys — OpenAI’s GPT Images 2.5 takes the top two spots on the Artificial Analysis Image Arena, generating higher quality images in less than half the time, while keeping the same price as GPT Image 2 Nearly

Quota reset GLM Coding Plan

GLM Coding Plan usage limits update

@zixuanli_: Some users saw quota being consumed during the free window. That’s usually because a subagent is still on GLM-5.3. For fully zero quota usage, set all subagent models to GLM-5.3-Flash in Settings. But use models based on the task: GLM-5.3 s

observed

1 source
  • @zixuanli_ — Some users saw quota being consumed during the free window. That’s usually because a subagent is still on GLM-5.3. For fully zero quota usage, set all subagent models to GLM-5.3-Flash in Settings. But

Policy Codex

OpenAI to retire GPT-5.3-Codex-Spark next week

OpenAI’s Théophile Sautter says GPT-5.3-Codex-Spark will be retired next week, citing declining usage and better available models. No exact calendar date is stated.

observed Next week

1 source
  • @thsottiaux — Next week we’ll be retiring GPT-5.3-Codex-Spark. Can you believe we shipped a model named as such!! It's had a good run and was a lot of fun, but usage has been declining and we have significantly

Just in

Topic radar

All →

Quick takes

Frontier model deployments are increasingly embedding domain-specific data and computer use capabilities directly into specialized vertical workflows, spanning quantitative financial modeling to autonomous software verification.Rapid progress in agentic swarms and recursive model improvement is intensifying internal governance debates at frontier labs, prompting dedicated board appointments and heightened scrutiny around catastrophic safety risks.Cross-border distillation campaigns are escalating competitive friction as frontier labs encounter unauthorized utilization of proprietary conversational outputs to train competing open and closed models.Developer platforms are formalizing managed agent harnesses and unified evaluation interfaces to support long-running autonomous execution across terminals, browsers, and enterprise codebases.OpenAI is pairing GPT-6 Astra with domain-specific workflows in coding and finance, though high user demand has prompted infrastructure constraints that paused Pro subscription sign-ups.Provider offerings are shifting toward managed agent APIs and harness-driven orchestration, supporting end-user productivity gains alongside automated internal research workflows.Accelerating agentic capabilities and recursive self-improvement dynamics are intensifying internal safety debates and governance adjustments across frontier research laboratories.AI infrastructure expansion is navigating regional clean power restrictions alongside ongoing gas utilization, accompanied by substantial overseas data center capital commitments.

Just in

Market pulse

bullish
AI pulse 58/100

AI-linked equities are broadly positive, with Super Micro Computer +7.28%, Arm Holdings plc +4.17%, Advanced Micro Devices +2.49% leading the tracked basket.

TickerCompanyMove
SMCI Super Micro Computer +7.28%
ARM Arm Holdings plc +4.17%
AMD Advanced Micro Devices +2.49%
AMZN Amazon.com +1.94%

Landscape

Coding Agents

Model Releases

Hot Builder Skills

AI Infrastructure

Research & Evals

Recent leads

History →