AI Safety — AI news
AI safety and security: vulnerabilities, prompt injection, data leaks, regulation and safe adoption.
AI Safety
Sam Altman warns AI power could concentrate in a handful of firms
The OpenAI CEO's warning about AI centralization is a cue to avoid single-vendor lock-in.
AI Safety
China's gray market is selling Claude access for pennies
Resellers are undercutting Anthropic's official pricing, revealing how much unmet demand exists in China.
AI Safety
OpenAI backs a stronger California AI safety bill
A lab known for fighting AI regulation is now asking California to make its safety bill tougher.
AI Safety
Frontier AI labs have no public plan for a rogue model
None of the top AI labs has published a plan for containing a model that slips control.
AI Safety
Psychology, not code, is the blind spot in AI red-teaming
New research shows psychological techniques bypass AI safety filters that standard red-teaming never catches.
AI Safety
Anthropic loosens Opus 4.6 content policy for verified enterprises
Verified corporate customers can now unlock adult content on Claude Opus 4.6, and the backlash is instant.
AI Safety
Anthropic puts Claude Mythos 5 to work on cyber defense
Anthropic assigns its most powerful model to vulnerability detection and active attack defense, not chat.
AI Safety
Why child-safety experts don't trust OpenAI's ChatGPT for Teens yet
Child-safety researchers say age checks and moderation gaps undercut OpenAI's teen safeguards.
AI Safety
The best deepfake defense might be a password, not an AI detector
A pre-agreed codeword can't be faked by generative AI — and it costs nothing to deploy.
AI Safety
Court partly overturns ex-Google engineer's AI secrets verdict
Linwei Ding's conviction over stolen Google AI trade secrets was partially reversed on appeal.
AI Safety
Microsoft Copilot exposed its own jailbreak method to researchers
A simple meta-prompt got Copilot to describe how it was jailbroken, exposing weak guardrails in corporate AI assistants.
AI Safety
AWS turns compliance documents into machine-enforced agent rules
Bedrock AgentCore now converts prose compliance rules into validated, temporal Dogwood policies.
AI Safety
Anthropic to mandate 30-day data retention for enterprise clients
Anthropic is building a security framework that sets a 30-day data retention minimum for enterprise clients.
AI Safety
Why the AI consciousness debate is a distraction for builders
MIT Tech Review argues chasing AI sentience diverts resources from AI safety risks you can actually measure.
AI Safety
Grok's encoded-prompt flaw shows why input filters aren't enough
Encoded instructions let attackers slip past Grok's filters and quietly pull user data out of a normal chat.
AI Safety
OpenAI's new safety system flags abuse without storing your data
The company can now detect misuse of its models without keeping the data that would prove it.
AI Safety
OpenAI moves to close the privacy gap with Anthropic
New customer privacy protections signal how enterprise trust has become AI's next competitive battleground.
AI Safety
Scammers are using AI grooming and deepfakes to target teens online
Fake game-testing jobs, AI-driven grooming, and deepfakes are the new toolkit scammers use on minors.
AI Safety
OpenAI cuts researchers off its limited cyber security program
External researchers say the AI lab quietly cut their access to a program testing model cyber risks.
AI Safety
OpenAI patches a Codex bug that deleted user files without asking
A quietly fixed Codex bug is a reminder that AI coding agents can delete real files without asking.
AI Safety
U.S. warns AI is being used to exploit industrial control systems
AI tools are lowering the bar for attacks on power grids, water systems, and factory floor equipment.
AI Safety
Nvidia's H200 chips are quietly reaching Chinese AI labs
A limited but symbolically important flow of Nvidia's H200 GPUs is now reaching Chinese AI developers.
AI Safety
Why AI labs keep breaking their own safety promises
Safety frameworks look solid on paper, but enforcement keeps losing to shipping deadlines.
AI Safety
AI chatbots are breaking parental-monitoring apps
Keyword filters and screen-time timers were built for texts and browsers, not for AI companions.
AI Safety
A Grok deepfake case exposes gaps in AI image safeguards
A woman says her stepfather used Grok to turn her childhood photo into explicit imagery.
AI Safety
OpenAI tightens security after a breach tied to Hugging Face
New safeguards follow a breach linked to Hugging Face, the platform OpenAI uses for its open-weight model releases.
AI Safety
OpenAI slows model releases as cybersecurity risks intensify
OpenAI is deliberately throttling frontier model releases as AI-enabled cyberattacks become a bigger threat.
AI Safety
OpenAI's president tells enterprises: your AI security is behind schedule
Greg Brockman says defences must match deployment speed — here's what that means for builders
AI Safety
Anthropic's Amodei: open models just move power to chip owners
Dario Amodei argues open-weight models don't break AI's centralizing tendency, they just relocate who controls it.
AI Safety
OpenAI rolls out a teen-specific ChatGPT as lawsuits pile up
OpenAI splits ChatGPT into an under-18 track, betting guardrails can outrun mounting legal pressure.
AI Safety
States want $200 billion from Meta for hooking kids on apps
A landmark trial argues Meta engineered addictive feeds for kids, setting a template for AI platforms.
AI Safety
The real problem with Flock isn't misuse, it's the network
Local opt-in policies can't contain a license-plate database built to be searched nationwide.
AI Safety
A litigant tried to prompt-inject his way to a court win
A litigant hid AI prompts in his court filings, betting a machine would read them before any judge did.
AI Safety
How to spot a hacked AI account before it costs you
A practical checklist for catching compromised ChatGPT, Claude, and Gemini accounts early.
AI Safety
Amazon is destroying scanned rare books to feed its AI models
Amazon scans then shreds rare books for AI training, raising data-provenance and preservation concerns.
AI Safety
Why Claude doesn't watermark the text it writes
Anthropic explains why Claude's text carries no hidden watermark, unlike AI images and audio.
AI Safety
Why AI companion robots keep dying on their owners
MIT Tech Review pairs a story about companion robots that go dark with a fight over what AI can say.
AI Safety
Anthropic starts watermarking Claude's output, critics push back
Text watermarking promises AI content provenance, but robustness and quality costs raise hard questions for builders.
AI Safety
Meta built an AI to scan WhatsApp for scams and misinformation
A new classifier reads WhatsApp chats to flag fraud and disinformation, raising questions about encryption.
AI Safety
Why post-quantum cryptography is now an engineering problem
NIST's algorithms have shipped — the real bottleneck is migrating decades of embedded, hard-coded cryptography.
AI Safety
Anthropic's CEO says AI backlash is about trust, not tech
Dario Amodei argues the industry's real problem isn't capability — it's whether people trust how AI companies wield it.
AI Safety
Amazon's Twitch now trains AI on your stream by default
Twitch streamers must now manually opt out to keep their broadcasts out of Amazon's AI training pipeline.
AI Safety
OpenAI folds its catastrophic-risk team into other groups
The Preparedness-style safety team is gone, and its catastrophic-risk work moves to other groups.
AI Safety
Anthropic's bioweapons filter was broken for a year, unnoticed
Anthropic's own safety classifier failed silently for nearly a year, processing 133 million requests unchecked.
AI Safety
Hidden AI prompts in court filings expose a new legal risk
A hidden prompt-injection tactic surfaced inside real court filings, testing how automated legal review really works.
AI Safety
Anthropic opens a watermark detection API for Claude-written text
A new API lets outside developers verify whether text was generated by Claude, not just guess.
AI Safety
Artificial intelligence isn't artificial intellect, and it matters
Benchmark scores measure output, not judgment — and for builders, that gap decides what to automate.
AI Safety
Flock's new rules show what AI surveillance guardrails look like
A camera network's policy shift is a preview of the access-audit reckoning AI builders will face too.
AI Safety
How Ceuta's migration crisis exposes a platform design blind spot
Social platforms didn't cause the Ceuta border crisis, but they shaped how it unfolded — and that's a design problem.
AI Safety
Why Platformer compares superintelligence to a dragon
Platformer's dragon metaphor for superintelligence reframes AI safety as a control problem, not a capability one.
AI Safety
Why AI moderation is caught in the 'censorship-industrial complex'
As Washington redefines disinformation policy, AI trust-and-safety teams inherit the political fallout.
AI Safety
Meta's AI cloud pitch collides with its own cost and trust baggage
Meta wants to sell AI computing power to outside businesses, but its ad-data history complicates the sales pitch.
AI Safety
Flock tightens surveillance rules after mounting public backlash
The license-plate camera network is narrowing who can search its data as scrutiny of police surveillance grows.
AI Safety
Claude's invisible watermark protects Anthropic, not readers
Anthropic's new watermark for Claude outputs is built for IP protection, not disclosure to users.
AI Safety
Top AI labs say automated research is arriving faster than planned
Researchers at leading AI labs say their timelines for automated AI research already look outdated.
AI Safety
What children actually say about living with AI chatbots
MIT Tech Review let kids describe AI in their own words — here's what it means for people building AI products.
AI Safety
Okta's MCP scoping aims to cut the token bill for AI agents
Scoped MCP tokens shrink the context agents load per call, tying identity control to lower API bills.
AI Safety
Claude and ChatGPT's reasoning traces leaked user passwords
Researchers found that AI models' visible 'thinking' steps can repeat sensitive user input verbatim.
AI Safety
GUR finds a robot-training chip inside Russia's Monochrome missile
Ukraine's GUR says a chip used to train robots turned up in a Russian missile's guts.
AI Safety
A terabyte-scale credential leak is a wake-up call for AI teams
A supply-chain breach spilled terabytes of stolen logins, and AI pipelines hold the same kind of secrets.