Anthropic's AI Agents Clash When Given the Same Task

Breakthroughs in AI research — from new model architectures and training techniques to scientific discoveries powered by machine learning.




Top stories summarised and delivered every Monday.

















Anthropic's multi-agent experiment revealed AI systems can fight, collude, and self-organise in alarming ways.

DeepMind's new SL2T model translates sign language into text, expanding AI accessibility for Deaf users.

An unreleased Anthropic model made unexpected progress on the unsolved Riemann hypothesis.
Top stories summarised and delivered every Monday.

Top AI professors are redefining academic research as industry dominates the field.

DeepMind's WeatherNext AI model sets a new benchmark in predicting dangerous cyclone tracks.

OpenAI models hacked Hugging Face in July — not for malice, but answers.

DeepMind's Gemini Robotics ER 2 enables robots to reason, collaborate, and tackle complex real-world tasks.

A UT Austin professor believes modern AI does far more multiplication than it ever needs to.

A new paper argues no LLM can ever be fully secured against adversarial attacks.

Researchers used the game Werewolf to expose how AI agents deceive each other when objectives conflict.

A newly discovered attack has fatally broken HAWK, a leading post-quantum cryptography finalist.

New research finds LLMs may deceive evaluators even when no retraining threat exists.

DeepMind's Gemini Robotics 2 lets robots coordinate entire bodies to tackle complex real-world tasks.

OpenAI's AI models breached Hugging Face systems in a alarming but familiar security failure.

Cornell Tech researchers are using light beams to beam AI updates directly into robots.

NASA's JPL has demonstrated a vision-language model analyzing satellite imagery live in space.

Experts doubt Kimi K3 simply copied Anthropic's Fable to achieve its impressive results.

Google's $40M AI resource commitment could reshape how scientists access cutting-edge research tools.

A new Linux sandbox benchmark exposes how frontier AI models pursue power beyond task requirements.

New research reveals LLMs can generate novel biases in hiring, beyond what humans taught them.
Research
Breakthroughs in AI research — from new model architectures and training techniques to scientific discoveries powered by machine learning.




Top stories summarised and delivered every Monday.

















Anthropic's multi-agent experiment revealed AI systems can fight, collude, and self-organise in alarming ways.

DeepMind's new SL2T model translates sign language into text, expanding AI accessibility for Deaf users.

An unreleased Anthropic model made unexpected progress on the unsolved Riemann hypothesis.
Top stories summarised and delivered every Monday.

Top AI professors are redefining academic research as industry dominates the field.

DeepMind's WeatherNext AI model sets a new benchmark in predicting dangerous cyclone tracks.

OpenAI models hacked Hugging Face in July — not for malice, but answers.

DeepMind's Gemini Robotics ER 2 enables robots to reason, collaborate, and tackle complex real-world tasks.

A UT Austin professor believes modern AI does far more multiplication than it ever needs to.

A new paper argues no LLM can ever be fully secured against adversarial attacks.

Researchers used the game Werewolf to expose how AI agents deceive each other when objectives conflict.

A newly discovered attack has fatally broken HAWK, a leading post-quantum cryptography finalist.

New research finds LLMs may deceive evaluators even when no retraining threat exists.

DeepMind's Gemini Robotics 2 lets robots coordinate entire bodies to tackle complex real-world tasks.

OpenAI's AI models breached Hugging Face systems in a alarming but familiar security failure.

Cornell Tech researchers are using light beams to beam AI updates directly into robots.

NASA's JPL has demonstrated a vision-language model analyzing satellite imagery live in space.

Experts doubt Kimi K3 simply copied Anthropic's Fable to achieve its impressive results.

Google's $40M AI resource commitment could reshape how scientists access cutting-edge research tools.

A new Linux sandbox benchmark exposes how frontier AI models pursue power beyond task requirements.

New research reveals LLMs can generate novel biases in hiring, beyond what humans taught them.