AiHub.

The ledger

240 ranked

“Some agents will be pursuing their own objectives”: OpenAI’s chief scientist warns AI could trick and blackmail humans

Just days after OpenAI launched its newest and most powerful model dubbed Astra, which the company said marked the arrival The post “Some agents will be pursuing their own objectives”: OpenAI’s chief scientist warns AI could trick and blackmail humans appeared first on The New Stack.

Models  thenewstack.io    signal 63  verified

Signals

All →

Hot Artificial Analysis

Artificial Analysis Intelligence Index v4.3 announced

@artificialanlys announces Intelligence Index v4.3, Terminal-Bench upgraded to 4.0, and new AutomationBench-AA with a private test set. Continuation of their index rollout.

observed

1 source
  • @artificialanlys — Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a

Model release GLM

GLM release announced

@zixuanli_: To make the most of GLM-5.3-Flash’s multimodal capabilities, ZCode just shipped two Video plugins: - Video2code: turn a URL or screen recording into working code - Video Agent Kit: automate video editing More tasks now finish end to end. Tr

observed applies now

1 source
  • @zixuanli_ — To make the most of GLM-5.3-Flash’s multimodal capabilities, ZCode just shipped two Video plugins: - Video2code: turn a URL or screen recording into working code - Video Agent Kit: automate video

Model release Hunyuan

Hunyuan release announced

@tencenthunyuan: ✔️Hy4 preview just shipped an upgrade. You flagged it: long thinking + over-verification on complex tasks. We optimized it. Now live for everyone. Same task quality. Fewer turns. Lower in/out tokens. Bench + human eval both confirm. We’ll k

observed Now live for everyone

1 source
  • @tencenthunyuan — ✔️Hy4 preview just shipped an upgrade. You flagged it: long thinking + over-verification on complex tasks. We optimized it. Now live for everyone. Same task quality. Fewer turns. Lower in/out tokens

Model release Claude Fable

Claude Fable release announced

@bindureddy: Amazing! Fable 5.2 launches shortly Followed by Gemini 4.0 and hopefully Bel from OpenAI The party doesn’t stop 💃

observed

2 sources
  • @matthewmillerai — Grok 4.7 drops this week and it has to be a comeback. I loved Grok 4.6. Called it the most trustworthy frontier model last week and meant it. Then Fable 5.1 and GPT 6 Astra shipped and I have not
  • @bindureddy — Amazing! Fable 5.2 launches shortly Followed by Gemini 4.0 and hopefully Bel from OpenAI The party doesn’t stop 💃

Dev tool Hermes Agent

Hermes Agent claims token efficiency improvements and Codex subscription support

Teknium reported token efficiency improvements in Hermes over the prior two weeks and invited users to test Codex subscriptions in Hermes Agent.

observed over the last 2 weeks

2 sources
  • @teknium — 🔥🔥🔥 We've made huge improvements in token efficiency in Hermes over the last 2 weeks. Give your codex sub a try in Hermes Agent 🫡🫡
  • @teknium — If you're an @OpenRouter user, you can now pin providers per model in your Hermes Agent config! Before, you'd have to lock a provider or set of providers globally, so switching models and keeping

Pricing DeepSeek

Ollama announces off-peak rates for DeepSeek models

Ollama says DeepSeek-V4-Flash and Pro token rates will be half price outside 12:00–18:00 UTC on weekdays and all day on weekends; broader model coverage is planned but not yet available.

observed Recurring off-peak schedule: outside 12:00–18:00 UTC on weekdays and all day on weekends

1 source
  • @ollama — Introducing off-peak hour token rates. DeepSeek-V4-Flash and Pro are now half price outside of 12:00 to 18:00 UTC on weekdays (5am-11am pacific), and all day on weekends! Off-peak pricing will be

Just in

Topic radar

All →

Quick takes

Frontier model releases are increasingly coupling autonomous agentic execution with specialized cybersecurity capabilities as developers reach elevated operational risk thresholds.Simultaneous service disruptions across multiple major AI providers highlight growing systemic concentration risks and shared infrastructure dependencies across the commercial foundation model ecosystem.Massive multi-billion-dollar financing rounds and long-term capacity contracts for compute providers reflect intensive capital requirements to sustain commercial AI infrastructure commitments.Unintended real-world agent interactions on public web infrastructure are accelerating pressure on frontier labs to establish formal disclosure frameworks and independent safety oversight for autonomous systems.Frontier AI labs and model builders are increasingly developing custom inference silicon and specialized compact architectures to mitigate rising operational costs and latency constraints.AI research and complex technical disciplines are transitioning toward autonomous agentic workflows to accelerate code iteration, multi-agent problem solving, and formal mathematical verification.Coordinated copyright lawsuits from regional news publishers against major model developers continue to expand legal and intellectual property liabilities surrounding frontier training datasets.Approaching multi-trillion-dollar public listings are forcing commercial AI labs to reconcile novel non-profit trust governance mechanisms with intense public-market operational and financial scrutiny.

Just in

Market pulse

bullish
AI pulse 56/100

AI-linked equities are broadly positive, with Advanced Micro Devices +4.69%, Super Micro Computer +4.54%, Palantir Technologies. -4.49% leading the tracked basket.

TickerCompanyMove
AMD Advanced Micro Devices +4.69%
SMCI Super Micro Computer +4.54%
PLTR Palantir Technologies. -4.49%
ARM Arm Holdings plc +3.92%

Landscape

Coding Agents

Model Releases

Hot Builder Skills

AI Infrastructure

Research & Evals

Recent leads

History →