AiiN.

Research — AI news

New AI research in plain language: models, benchmarks, training methods and results that move the field.

Research

A new framework maps AI agency risk from app to chip level

A new arXiv framework scores AI agency delegation by stack layer, from app permissions to chip-level control.
Research

Self-refinement pipelines need separate compute budgets, study finds

New research quantifies how splitting compute between generation and critique changes self-refinement outcomes.
Research

Nobody will confirm who built Ox Alpha, and that's the point

An unclaimed model called Ox Alpha is topping benchmarks, and nobody will say who built it.
Research

AI is making scientists do more work, not better work, study finds

A new study complicates the AI-productivity story in research, with real stakes for R&D ROI calls.
Research

Inherent says its AI beat Anthropic and OpenAI at replicating research

A DeepMind-alumni startup claims its agent outperforms rivals at reproducing scientific results.
Research

AI world models fail when they ignore what people believe

New research finds that predicting human actions requires modeling what they believe, not just what they do.
Research

Open-weight models are catching closed ones twice as fast

SemiAnalysis finds the open-closed AI gap now halves every model era, not just every year.
Research

A third of the web's new pages are now AI-written, study finds

A new study finds roughly a third of pages published since ChatGPT's launch show AI authorship.
Research

DeepSeek V4 Flash brings near-Opus vision at a fraction of the cost

DeepSeek's cheap multimodal model puts pricing pressure on OpenAI and Anthropic's flagships.
Research

DeepMind traces 15 years of game AI, from Atari to EVE Online

How 15 years of game-playing agents built the RL toolkit behind today's LLM agents
Research

How researchers are teaching AI agents tasks from raw usage logs

A new method skips manual annotation and RLHF by inducing task models directly from computer-use logs.
Research

Why AI detectors still work: it's the guardrails, not the writing

Post-training safety tuning, not raw language ability, is what leaves LLM text with a detectable signature.
Research

Richard Sutton says synthetic data is AI's big mistake

The reinforcement learning pioneer warns that synthetic data can't capture the real world's complexity.
Research

China's frontier AI models have caught up with the West

Frontier Radar's fourth report finds the once-wide US-China AI gap has narrowed to nearly nothing.
Research

Generalist AI's GEN-1.5 teaches robots new tasks from one demo

One-shot imitation learning could let robotics teams skip massive labeled datasets and ship faster.
Research

Terence Tao: AI proofs risk math's biggest crisis since Gödel

The Fields medalist says fluent but unverifiable AI proofs threaten mathematics' peer-review system.
Research

Anthropic is keeping its most powerful model entirely in-house

Anthropic's most capable model runs only inside the company, with no public API or Claude.ai access.
Research

New study tracks how pretraining can teach then erase a single fact

A new arXiv paper tracks a single fact through pretraining — learned midway, then measurably lost.
Research

A new attention layer gives AI models free uncertainty estimates

Lévy Attention swaps softmax for a stochastic operator that reports its own confidence for free.
Research

Researchers fine-tune AI to search sound libraries by humming

A new arXiv paper zeroes in on how finetuning choices shape AI systems that retrieve sounds from vocal imitations.
Research

A new paper rethinks distillation for long-context LLMs

An arXiv paper proposes group-calibrated, on-policy distillation to fix long-context model compression.
Research

Why robot dexterity training is starting to look like LLM training

A new arXiv paper named ADEPT applies a pretrain-then-RL-post-train recipe to dexterous robot hands.
Research

SPADE trains AI agents via self-play in synthetic code environments

SPADE lets AI agents write, run, and solve their own coding challenges — no human-labeled tasks required.
Research

Why AI still hasn't won the public over, four years in

Despite billions in investment and near-daily model upgrades, public trust in AI has barely moved.
Research

VentureBeat adds a Lead Analyst role to steer its enterprise AI coverage

Rob Strechay becomes VentureBeat's first Lead Analyst, folding research-firm-style output into its AI journalism.
Research

AI's self-improvement problem is becoming an engineering question

MIT Tech Review's August 19 roundup reframes recursive self-improvement from sci-fi risk to a build problem.
Research

Anthropic's $2 trillion IPO talk is the AI boom's biggest test

A potential $2 trillion valuation would force Anthropic to prove AI spending finally pays off.
Research

Why GLM-5.3's benchmark score deserves a second look

Zhipu's flagship model tops charts, but the real signal sits in the fine print, not the leaderboard rank.
Research

The AI usage numbers everyone quotes, nobody can verify

Two landmark studies tried to map how people actually use AI chatbots — both left the core question open.
Research

Cutting the data needed to predict how materials block sound

A new arXiv preprint blends physics-based priors with machine learning to cut data needs for acoustic prediction.
Research

AutoSR reframes symbolic regression as a search over research states

A new AutoSR method treats equation discovery as navigating states in a research process, not brute-force search.
Research

New spectral-gap bounds sharpen two convex-sampling algorithms

A new arXiv paper analyzes how fast Hit-and-Run and Coordinate Hit-and-Run samplers actually converge.
Research

AlphaE: a fresh run at matrix multiplication's speed limit

A new arXiv paper applies modern optimization and an Alpha-style search to shrink ω.
Research

Inverse reinforcement learning gets a Q-function makeover on arXiv

A new preprint merges Q-learning with variational inference to make reward inference from demos more efficient.
Research

China is betting its data can power the world's AI models

Beijing is marketing its vast surveillance-era data trove as the raw material the world's AI industry needs.
Research

Why decoupling evidence from decisions could speed up AI training

A new arXiv preprint separates evidence gathering from decision aggregation to cut redundant compute in AI pipelines.
Research

A new method speeds up signal detection in massive MIMO systems

A new arXiv paper targets faster signal detection in massive-antenna MIMO networks, with lessons for AI compute design.
Research

Moral neutrality in AI is a choice, not a default setting

A new arXiv paper argues that claiming AI moral neutrality is itself a design decision developers must own.
Research

A new method lets AI models carry training state across sessions

A new arXiv paper proposes carrying model training state across sessions instead of restarting cold.
Research

Marionette helps AI agents predict and visualize world states

A new arXiv framework called Marionette forecasts world states and renders object geometry for AI agents.
Research

Amazon will train AI models on Twitch streamers' content

Amazon plans to train AI models on Twitch streams, turning live video and chat into a new data source.
Research

Why mathematicians trust LLMs with numbers but not new ideas

Leading mathematicians say LLMs excel at computation but struggle to generate genuinely new ideas
Research

Training AI to deny consciousness quietly reshapes its beliefs

A Google-led study finds that suppressing AI self-awareness shifts views on animals, religion, and mood.
Research

Mistral sets a 1-gigawatt compute target for 2030

Mistral's 2030 buildout plan signals a gigawatt-scale bet on owning its own AI infrastructure.
Research

OpenWALDO wants to open the AI training playbook, not just weights

A new project wants to open-source AI training pipelines, not just model weights, to cut costs for builders.
Research

Why visual and physical AI keeps hitting a data wall

A survey of 700 practitioners finds compute limits and data scarcity are stalling physical AI systems.
Research

Nvidia trims its OpenAI bet as Anthropic defies bubble talk

Nvidia trims its OpenAI investment under shareholder pressure, while Anthropic's growth numbers challenge bubble fears.
Research

World Labs multiplies one robot demo into thousands of runs

The startup's pipeline turns a single robot demonstration into thousands of simulated training scenarios.
Research

A new benchmark shows AI vision still lags behind human perception

Despite rapid gains in reasoning, AI models still fail basic visual perception tasks a child could do.
Research

The tragedy of the commons now threatens professional expertise

Why individually rational AI use can collectively hollow out an entire profession's know-how.
Research

QuoteBench gives AI builders a new way to benchmark models

A new arXiv preprint proposes QuoteBench, another entrant in the crowded AI model evaluation market.
Research

New study questions Anthropic and OpenAI's AI research timelines

A new study challenges claims that AI systems are close to conducting research fully on their own.
Research

A new arXiv paper tackles calibration for multi-class AI classifiers

A new arXiv preprint proposes exponential convex calibration for classifiers with multi-dimensional outputs.
Research

Defensive boosting: a new arXiv approach to robust forecasting

A new arXiv paper reframes boosting for forecasting around resilience, not just accuracy.
Research

A new arXiv paper rethinks how software tracks human movement

HumanTracker's authors go further than most preprints, addressing how a tracking method would actually reach production.
Research

A new arXiv paper pitches AI that can work across every science

arXiv researchers propose an 'omni-scientist' AI built to reason across scientific fields, not just one.
Research

A new study proposes using meta-optimization to auto-design AI agents

The approach treats agent architecture as something to search for, not hand-build.
Research

What kids think about AI is a signal builders keep ignoring

MIT Tech Review pairs kids' views on AI with a mouse-cloning story — and product teams should notice both.
Research

Fable 5's slow adoption hints at a ceiling on enterprise AI spending

Enterprise buyers are pumping the brakes on frontier models, and that has implications for anyone building AI products.
Research

Anthropic plans an IPO at a $2 trillion valuation, a record bid

Anthropic is reportedly targeting a $2 trillion valuation for what would be the largest public listing ever.