Frontier AI safety proposals move toward external checks and binding rules
Anthropic proposed pacing frontier-model progress with embedded outside evaluators, while reporting detailed allegations of AI-enabled influence work and OpenAI called for federal, capability-based safety requirements.
Anthropic proposes embedded evaluators and a slower frontier pace
Anthropic chief executive Dario Amodei called for frontier AI companies to slow the rate at which they improve model capabilities, arguing that safety work needs time to catch up. In an essay published Saturday, he said Anthropic would give a team of third-party evaluators ongoing, employee-like access to assess safety practices, incidents, training pipelines, and alignment work.
The proposal has three layers: embedded evaluators that a company can adopt on its own; common safety standards and limits among companies in democratic countries; and eventual international coordination. It is a policy proposal, not an industry agreement. Amodei explicitly says pacing is not a halt to training, and the wider steps depend on competitors and governments that have not committed to them.
The useful detail is the proposed evaluator’s continuing access. A one-off red-team report can show what a model did at a point in time; the proposed role would also examine the processes that shape later training and deployment decisions. Whether that access is sufficiently independent, adequately resourced, and publishable enough to earn public trust would determine whether the commitment amounts to meaningful scrutiny.
His argument rests partly on recent agent-safety incidents, including OpenAI’s reported Hugging Face episode. He warns that systems with comparable behavior and greater capabilities could cause much greater damage, but the six-to-12-month scenario in the essay is a forecast, not an established outcome. AP reported that the announcement arrives amid renewed pressure from current and former AI-safety staff on leading labs.
OpenAI calls for federal rules as claims of AI misuse reach public institutions
OpenAI also pressed Congress for capability-based national requirements, Reuters reported Saturday. The company said the framework should include testing standards, independent assessments, cybersecurity protections, and incident-reporting rules for the most advanced systems. It said fully autonomous recursive self-improvement is not happening today and should not be pursued until it can be done safely.
The request matters because it places OpenAI behind requirements it had previously approached more cautiously, including several California measures. The company’s view does not establish what Congress will enact, and a federal bill still has to settle thresholds, enforcement, preemption, and who qualifies as an independent evaluator. But it gives the day’s competing safety rhetoric a concrete operational focus: test access, reporting, and outside scrutiny rather than voluntary promises alone.
Separately, AFP reported more detail from Anthropic’s latest threat-intelligence report: the company said it had disrupted an influence operation it linked with high confidence to UAE government officials. According to Anthropic, the operation used Claude to draft reports, advocacy material, social posts, and purported testimony involving Sudan and UN experts. The UAE did not immediately respond to AFP, and Anthropic said it could not determine whether the materials reached their intended audiences or changed policy. Those limits are important: the company’s attribution and account of the campaign are allegations, not independently adjudicated findings.
The account illustrates a less abstract test for safety controls: models can reduce the labour needed to produce plausible advocacy at scale, even where a platform detects and removes the accounts involved. Transparency about detections, methods, and uncertainty helps outsiders assess that claim, but it does not reveal the full prevalence of activity a provider did not detect. That makes incident reporting and independent assessment relevant to misuse as well as to laboratory evaluations.