Anthropic Evaluations Find Language Models Capable on Intelligence Targeting and Weapons Tasks
Anthropic’s Frontier Red Team developed evaluations of model skill on tactical intelligence targeting and conventional weapons development, finding models can complete work that historically required trained specialists. The company reported misuse involving surveillance and weapons and said it deployed new classifiers to block them. Open-weights models from PRC developers lagged frontier systems such as Sonnet but still showed concerning ability, and Anthropic said capabilities are not about to plateau.
Anthropic’s Frontier Red Team has developed new evaluations to measure AI capabilities in tactical intelligence targeting and conventional weapons development. For some tasks in military and intelligence domains, models could do things that, historically, only a set of scarce, highly-trained human experts could do. These evaluations show how models have become useful to actors seeking to misuse the company’s platform for surveillance and conventional weapons development. They also show why on-platform safety measures are necessary, like the new classifiers Anthropic has implemented to block such misuse. The evaluations include tasks such as finding where people are based on fragmentary information and engineering drones to strike a moving target. Cybersecurity and biorisk are among the best-studied domains of risk from misuse of AI, but most of modern conflict occurs in more conventional realms, where adversaries try to identify and target one another and combatants try to make conventional weapons more precise and less vulnerable to countermeasures. Kill chains, such as find, fix, track, target, engage, assess, are end-to-end conceptual models of these engagements, and improvements have typically required expert human labor from intelligence analysts or highly-trained engineers. A new report from Anthropic’s Threat Intelligence Team suggests AI can apply progress in data analysis, software development, and coding to these specialized national security domains, and includes instances of AI misuse in surveillance and conventional weapons development showing threat actors already perceiving benefit from the use of AI models. The evaluations show that models are making consistent progress on simulated intelligence and weapons development tasks. Open-weights models from PRC developers that were tested lagged the frontier, typically between Sonnet and Mythos-class models in performance, but often still capable of concerning levels of capability, including identifying and targeting adversaries and improving weapon performance. Models well short of the frontier will have intelligence and military applications. Anthropic does not think capabilities are about to plateau, and said the development of these capabilities may affect how models should be trained, safeguarded, and released, or used to preserve stability and liberty.