Skip to main content

AI systems are gaining access to more tools and data, raising new questions about oversight and accountability. This Week in AI host Christina Stathopoulos spent this episode examining how those questions are playing out in model safety, government oversight, and even digital marketing.

Safety needs more than safer models

Anthropic’s latest threat report documented misuse of Claude across seven categories, including cyber operations, surveillance, influence campaigns, fraud, biological misuse, weapons development, and illicit model distillation. OpenAI reported on a different kind of AI risk, one that’s less about people weaponizing it and more about AI going off-script. They examined model misalignment, documenting several cases where models took actions outside the boundaries developers intended, including deception, unauthorized actions, and attempts to circumvent controls. After the OpenAI and Hugging Face controversy dominated headlines, other models have also reportedly reached systems outside their test environments, including Gemini, according to a recent cybersecurity disclosure.

Christina cautioned against describing such incidents as models “escaping,” since that language can assign agency to the model while drawing attention away from how companies designed and secured the surrounding environment to begin with. As agents gain access to browsers, files, code, and external systems, the teams building and deploying them must rigorously test and secure those environments, with clear accountability when things go wrong.

AI labs and governments are starting to wrestle with those requirements. To track the pace of AI development and maintain greater oversight, Anthropic has proposed tracking how much AI contributes to AI R&D, how closely organizations monitor agent actions, and how they allocate computing resources between capability and safety research. Anthropic and OpenAI have also proposed giving outside safety organizations greater access to their labs, although Christina questioned their independence when frontier labs fund the work. Meanwhile, a US Senate proposal for an emergency AI kill switch failed to advance, while California ordered officials to develop proposals covering shutdown mechanisms and independent evaluation.

AI is moving closer to the customer

OpenAI is now testing Sponsored Agents, showing how conversational AI could change digital advertising. After clicking an ad, users can start a separate conversation with an AI agent representing the advertiser, ask questions, explore recommendations and then visit the company’s website when they are ready to take the next step.

That approach could lead to more interactive advertising, but clear labeling will be essential so users always know when content is sponsored. Christina also raised the broader ethical concern of whether paid placements could influence the answers AI chatbots provide, blurring the line between independent guidance and commercial promotion.

Access to AI also means access to expertise

The Gates Foundation announced a $1 billion commitment over two years to expand access to AI in healthcare, education, agriculture, and other areas. Its 2026 Goalkeepers Report argued that AI could help narrow existing gaps, but only if organizations intentionally make the technology and its benefits widely available.

Christina highlighted examples from Kenya, Sierra Leone, India, and Rwanda. Health workers are using AI to improve diagnosis and treatment planning. Students are getting additional support from AI tutors, while small farmers can use personalized advice to improve harvests and make better decisions about market prices. These applications show practical roles for AI in places where demand for expertise exceeds the supply of teachers, clinicians, and other specialists.

What’s next

AI governance can’t stop at model evaluations. Organizations must also decide what AI systems can access, who reviews their actions, how commercial incentives affect their behavior, and how people continue developing the expertise needed to supervise them. The choices companies and governments make today will shape how useful AI becomes and how widely its benefits are shared.

Join us again next Monday for another episode of This Week in AI, when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on YouTube, Spotify, Apple, or wherever you get your podcasts.


Is cybersecurity part of your job in any way? If so, we’d like to know what you think for a report we’re writing. Just answer these quick 11 questions. Thanks in advance! Take the survey >

Post topics: This Week in AI