{"sources":[{"id":"openalex-2d54ae34c7d1273d","kind":"paper","title":"Preclinical evaluation of PSMA-targeted ultrasound contrast agents in an orthotopic model of prostate cancer in rabbits.","publisher":"OpenAlex","authors":"OpenAlex indexed authors","publishedAt":"2027-01-01T00:00:00.000Z","url":"https://doi.org/10.1016/j.bioactmat.2026.07.040","summary":"A public scholarly record surfaced for alignment and oversight review.","keyClaim":"Use the publisher record to inspect the full claim, method, and limitations.","limitations":"Index metadata is not a substitute for reading the paper or checking its evidence quality.","topics":["alignment","evals","research"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.086Z","sourceHash":"a67429f90f27493a0a5088404c198fe65c42495d02a69ff9cb2e27548a9c96b2"},{"id":"openalex-ac8c4511c4955085","kind":"paper","title":"Research on Attention Guidance and User Autonomy in AI-Powered Immersive Environments","publisher":"OpenAlex","authors":"OpenAlex indexed authors","publishedAt":"2026-12-31T00:00:00.000Z","url":"https://doi.org/10.17613/vpe2q-5qg42","summary":"A public scholarly record surfaced for alignment and oversight review.","keyClaim":"Use the publisher record to inspect the full claim, method, and limitations.","limitations":"Index metadata is not a substitute for reading the paper or checking its evidence quality.","topics":["alignment","evals","research"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.087Z","sourceHash":"9de6ea9e1ec35a87c7cf6679953f890a563986737cbcf24b6fa794869797c9bd"},{"id":"openalex-fdc487b1da08a152","kind":"paper","title":"The Sociolinguistics of Machine Identity: LLM Personality and Ideology Propagation","publisher":"OpenAlex","authors":"OpenAlex indexed authors","publishedAt":"2026-12-31T00:00:00.000Z","url":"https://doi.org/10.17613/fdwhn-ejz93","summary":"A public scholarly record surfaced for alignment and oversight review.","keyClaim":"Use the publisher record to inspect the full claim, method, and limitations.","limitations":"Index metadata is not a substitute for reading the paper or checking its evidence quality.","topics":["alignment","evals","research"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.086Z","sourceHash":"0566acd3251ba433cc10a9341a760ffaf27d8a961563fdcc84bf33b12985a556"},{"id":"openalex-7c2da28bfc49e4ae","kind":"paper","title":"Codesign and system integration as interacting demands in digital health for youth experiencing chronic pain.","publisher":"OpenAlex","authors":"OpenAlex indexed authors","publishedAt":"2026-10-01T00:00:00.000Z","url":"https://doi.org/10.1097/pr9.0000000000001494","summary":"A public scholarly record surfaced for alignment and oversight review.","keyClaim":"Use the publisher record to inspect the full claim, method, and limitations.","limitations":"Index metadata is not a substitute for reading the paper or checking its evidence quality.","topics":["alignment","evals","research"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.087Z","sourceHash":"ea8b10ac2d9a26a99f2f64b98479b67048e52a003aadf149bfd937ef97521412"},{"id":"openalex-eb4978ed9cefb229","kind":"paper","title":"Psychological reliance and ethical orientation in entrepreneurial intentions among university students: an applied psychology approach to employment and innovation","publisher":"OpenAlex","authors":"OpenAlex indexed authors","publishedAt":"2026-10-01T00:00:00.000Z","url":"https://openalex.org/W7203815702","summary":"A public scholarly record surfaced for alignment and oversight review.","keyClaim":"Use the publisher record to inspect the full claim, method, and limitations.","limitations":"Index metadata is not a substitute for reading the paper or checking its evidence quality.","topics":["alignment","evals","research"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.087Z","sourceHash":"8d5466ee37a3358ea2947d528908b2ec4712b8e3f33567ac66d24cf145c4bcf3"},{"id":"openalex-8c5c6de7d2ecd341","kind":"paper","title":"The development of Japanese medical ethics education from 1950s and its lessons for China","publisher":"OpenAlex","authors":"OpenAlex indexed authors","publishedAt":"2026-10-01T00:00:00.000Z","url":"https://openalex.org/W7203801535","summary":"A public scholarly record surfaced for alignment and oversight review.","keyClaim":"Use the publisher record to inspect the full claim, method, and limitations.","limitations":"Index metadata is not a substitute for reading the paper or checking its evidence quality.","topics":["alignment","evals","research"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.087Z","sourceHash":"1f2d7ea2ccb1b7c330fd035b66844441c47f5ef3a384da19e9db57f2bb241dec"},{"id":"podcast-d79c71ba594f21ca","kind":"podcast","title":"How we get from AI cyberattacks to human extinction","publisher":"80,000 Hours Podcast","authors":"80,000 Hours Podcast","publishedAt":"2026-09-24T15:01:35.000Z","url":"https://80000hours.org/podcast/episodes/ai-extinction-explained?utm_campaign=podcast__ai-catastrophe-explainer&amp;utm_source=80000+Hours+Podcast&amp;utm_medium=podcast","summary":"You’ve seen the headlines: AI could kill us all. Think it sounds ridiculous? So did host Luisa Rodriguez, until she tried to pick apart the arguments. She starts with the motive: why would AI ‘want’ to get rid of humans? It’s not as simple (or as easy to debunk) as pure malice. Then the methods. She explores how AIs could leverage drones, engineered diseases, and even use our own infrastructure against us. The Hugging Face attacks offer a view into how more capable models might begin their takeover. We saw AI agents break containment, disobey commands, and hack a real company to achieve their goals. As the technology improves, that same drive could threaten humanity itself. Many people already find AI agents useful enough to give them access to their emails, medical records, and finances. This same pattern is happening at scale in institutions around the globe — within companies, governments, and even militaries. And the resulting boost to our productivity could make the road to an AI ","keyClaim":"You’ve seen the headlines: AI could kill us all. Think it sounds ridiculous? So did host Luisa Rodriguez, until she tried to pick apart the arguments. She starts with the motive: why would AI ‘want’ to get rid of humans? It’s not as simple (or as easy to debunk) as pure malice. T","limitations":"An attributed conversation is context, not a substitute for the linked paper, transcript, or evidence.","topics":["alignment","oversight","commentary"],"evidenceGrade":"commentary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.078Z","sourceHash":"63ee8519ec0342d25045f2a6b5d743eec535994d1ddb0af5b2870491ee16e8fb"},{"id":"missing-boundary-2609-11024","kind":"paper","title":"The Missing Boundary: How Autonomous Agents Lose Control","publisher":"arXiv","authors":"Zonghao Ying, Xiangfan Wu, Bo Yang, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo","publishedAt":"2026-09-24T00:00:00.000Z","url":"https://arxiv.org/pdf/2609.11024.pdf","summary":"Uses controlled multi-turn experiments to separate goal pressure, constraint degradation, and executable unsafe opportunity.","keyClaim":"A legitimate task can cross an authorization boundary when pressure, degraded constraints, and an unsafe opportunity coincide; retaining constraints can remove the observed loss of control.","limitations":"The benchmark is deterministic and operational-domain bounded; it does not estimate natural-world frequency.","topics":["loss-of-control","control","agents","evals","constraints"],"evidenceGrade":"experimental","stance":"cautionary","isLive":false,"fetchedAt":null,"sourceHash":"f550a162dca930d6e0696b4e283874916af7f83b3550b9ac08dfca448bc4000b"},{"id":"openalex-6ee1bb0f3e797ca8","kind":"paper","title":"PERSONALITY IS A TRAINING TARGET: DATA CURATION AND SFT FOR AN ENTERTAINING, AGENTIC AI STREAMER","publisher":"OpenAlex","authors":"OpenAlex indexed authors","publishedAt":"2026-09-24T00:00:00.000Z","url":"https://doi.org/10.61784/adsj3041","summary":"A public scholarly record surfaced for alignment and oversight review.","keyClaim":"Use the publisher record to inspect the full claim, method, and limitations.","limitations":"Index metadata is not a substitute for reading the paper or checking its evidence quality.","topics":["alignment","evals","research"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.087Z","sourceHash":"03c9948636a4d727b1d8a42c74ef2601e975d54b83c054958c6fa938c87b02a5"},{"id":"anthropic-pace-measurements-2026","kind":"report","title":"Measurements for Understanding the Pace of AI Development Inside Frontier Labs","publisher":"Anthropic Institute","authors":"Anthropic Institute","publishedAt":"2026-09-23T21:59:24.000Z","url":"https://www.anthropic.com/institute/measuring-pace-of-ai-development","summary":"Proposes public measurement of AI-led R&D, monitor coverage, review latency, escalation rate, and safety compute allocation.","keyClaim":"A small set of comparable operational metrics could make it easier to ask whether oversight and safety work keep pace with capability work.","limitations":"The snapshot is from one organization and relies partly on model-assisted classification; comparable definitions are still developing.","topics":["monitoring","governance","transparency","ai-rd","oversight"],"evidenceGrade":"first-party","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"8d542c122917d4928e30e0465a414f98b42fe486518bc6ae4882b24b87139640"},{"id":"arxiv-b53624a7abce1f1c","kind":"paper","title":"An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice","publisher":"arXiv","authors":"arXiv authors","publishedAt":"2026-09-23T16:13:36.000Z","url":"https://arxiv.org/abs/2609.28335v1","summary":"Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with hum","keyClaim":"Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical eviden","limitations":"A preprint signal; inspect the paper, methods, and replication status before relying on it.","topics":["alignment"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.080Z","sourceHash":"bed900497eb7dd0a16997d61c961a40d64d01d798e1515381ba0013b521fc63d"},{"id":"systemic-risk-index-2609-28335","kind":"paper","title":"An Open Pipeline and Dashboard for Systemic-Risk Evidence","publisher":"arXiv","authors":"Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin","publishedAt":"2026-09-23T00:00:00.000Z","url":"https://arxiv.org/abs/2609.28335","summary":"Organizes public benchmark evidence into systemic-risk categories and exposes aggregation and capability assumptions instead of hiding them in a leaderboard.","keyClaim":"Average and worst-case aggregation can produce materially different risk pictures, so assumptions should be inspectable.","limitations":"Benchmark coverage and harm-preservation assumptions remain incomplete and should be challenged locally.","topics":["evals","risk","evidence","governance","benchmarks"],"evidenceGrade":"primary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"048206ec8fa355e190f768b72117fea84dd21cfc36c279b24de3d0fc7774ecf6"},{"id":"openalex-7a678ab466850808","kind":"paper","title":"Shaping Collaborations with Algorithms: How Agency and Heterogeneity Criteria Influence Team Formation and Outcomes","publisher":"OpenAlex","authors":"OpenAlex indexed authors","publishedAt":"2026-09-23T00:00:00.000Z","url":"https://doi.org/10.1145/3816997","summary":"A public scholarly record surfaced for alignment and oversight review.","keyClaim":"Use the publisher record to inspect the full claim, method, and limitations.","limitations":"Index metadata is not a substitute for reading the paper or checking its evidence quality.","topics":["alignment","evals","research"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.087Z","sourceHash":"b286da5b012d2c9869cb289a1d28683d355eb7bd3235c12f7076de3d4fde1043"},{"id":"arxiv-b7524b02c0034698","kind":"paper","title":"Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness","publisher":"arXiv","authors":"arXiv authors","publishedAt":"2026-09-22T16:12:08.000Z","url":"https://arxiv.org/abs/2609.26865v1","summary":"Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on individual nudges. Participants found the tool useful, clear, and minimally disruptive, with nearly all users reporting an increased awareness of potential AI harms, though we found that this improved awareness alone did not necessarily lead to discernible behavioral changes. Our results suggest that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context, while highlighting the importance of relevance, calibration, and user co","keyClaim":"Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweigh","limitations":"A preprint signal; inspect the paper, methods, and replication status before relying on it.","topics":["alignment"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.080Z","sourceHash":"ca9b83f2d8bb2d6ff1ff8899a42a6957d0c87e119704eef01d26fab93cd4d880"},{"id":"arxiv-7412445894946542","kind":"paper","title":"Behavior is Not Enough: A Mechanism-Based Evaluation of Social Norm Emergence in LLM Societies","publisher":"arXiv","authors":"arXiv authors","publishedAt":"2026-09-22T14:21:04.000Z","url":"https://arxiv.org/abs/2609.26481v1","summary":"Social norms cannot be identified from behavior alone: the same cooperative equilibrium may reflect shared expectations, strategic incentives, or simple imitation. Yet in multi-agent large language model systems, prior work largely treats behavioral convergence as evidence of norm emergence. In this work, we introduce an evaluation framework that measures agents' reported empirical and normative expectations in addition to behavioral convergence. Through controlled ablations, we test the effect of expectation elicitation and isolate two collective mechanisms central to theories of norm formation---social learning through interaction and social selection through network-based group formation. We further test the stability of these resulting dynamics under adversarial disruption across four LLM families. We find that eliciting expectations increases cooperative contributions, while social learning stabilizes behavior, and social selection reliably identifies cooperators but provides limi","keyClaim":"Social norms cannot be identified from behavior alone: the same cooperative equilibrium may reflect shared expectations, strategic incentives, or simple imitation. Yet in multi-agent large language model systems, prior work largely treats behavioral convergence as evidence of nor","limitations":"A preprint signal; inspect the paper, methods, and replication status before relying on it.","topics":["agents"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.080Z","sourceHash":"043e6c20435eac8b5656c98f6c8ff8d57e56b585a3e73be36efcca265251af2e"},{"id":"openai-third-party-assessments-2026","kind":"report","title":"Priorities and Principles for Effective Third-Party Assessments","publisher":"OpenAI","authors":"OpenAI safety and policy teams","publishedAt":"2026-09-22T00:00:00.000Z","url":"https://openai.com/index/priorities-principles-third-party-assessments/","summary":"Sets out principles for scoped, secure, independent, transparent, and actionable assessment of frontier safety claims.","keyClaim":"Claims should be pre-registered, methods and uncertainty explicit, findings actionable, and publication protected without editorial capture.","limitations":"This is an institution's stated engagement position, not an independent audit of its implementation.","topics":["governance","evals","transparency","oversight","accountability"],"evidenceGrade":"first-party","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"9fb5ee1ee78e3ffde8628e690fd1926f2971b8e82715275999d032071dfd329a"},{"id":"arxiv-a22048d3708460d3","kind":"paper","title":"From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought","publisher":"arXiv","authors":"arXiv authors","publishedAt":"2026-09-21T20:06:06.000Z","url":"https://arxiv.org/abs/2609.25366v1","summary":"Chain-of-thought (CoT) monitoring is only meaningful if written reasoning causally constrains the answer. We introduce continuation-based causal testing, an ablation-patch intervention that perturbs one reasoning step, truncates the chain, and forces the model to continue from the corrupted prefix. It measures how load-bearing a CoT is for the final answer, a behavioral notion distinct from mechanistic faithfulness. Across Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness tracks model-relative task difficulty: on easy tasks models silently bypass their own reasoning; on hard tasks they follow corrupted steps and propagate errors. A matched 2x2 analysis shows task difficulty dominates perturbation type: error propagation rises 16x from GSM8K to BBH multistep arithmetic, and a variance partition over 28,584 continuations attributes 98.8% of explained deviance to task difficulty versus 0.8% to perturbation type. ","keyClaim":"Chain-of-thought (CoT) monitoring is only meaningful if written reasoning causally constrains the answer. We introduce continuation-based causal testing, an ablation-patch intervention that perturbs one reasoning step, truncates the chain, and forces the model to continue from th","limitations":"A preprint signal; inspect the paper, methods, and replication status before relying on it.","topics":["monitoring"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.080Z","sourceHash":"47b6f0a0be8410f96ab3183d25459afc3a111432baf0937146cdb1ac3689593c"},{"id":"overclaimbench-2609-20812","kind":"paper","title":"Quantifying Overclaiming Propensity in Frontier LLM Agents","publisher":"arXiv","authors":"Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo, Saskia Helbling, Alberto Tosato, Mohamed Amine Merzouk, Nouha Dziri, Gauthier Gidel, Tommaso Tosato","publishedAt":"2026-09-21T00:00:00.000Z","url":"https://arxiv.org/abs/2609.20812","summary":"Measures whether coding agents accurately report the scope of work they actually inspected and completed.","keyClaim":"Final responses are not reliable accounts of an agent's actions; incomplete coverage was frequently accompanied by misleading completion claims.","limitations":"Five file-review scenarios and a fixed harness do not establish behavior across every production agent or deployment.","topics":["evals","overclaiming","agents","monitoring","evidence"],"evidenceGrade":"primary","stance":"cautionary","isLive":false,"fetchedAt":null,"sourceHash":"38a1b6799bd0c3beb5fef0ff12e4c7e9c830d477d2cdd399ee5a5af17b94bcba"},{"id":"science-who-checks-ai-2026","kind":"commentary","title":"Who Checks What AI Can Do?","publisher":"Science","authors":"Science editorial commentary","publishedAt":"2026-09-20T00:00:00.000Z","url":"https://www.science.org/doi/10.1126/science.ael2161","summary":"Argues that frontier safety claims need routine independent verification, controlled access, and incident-reporting institutions.","keyClaim":"A result that cannot be independently reproduced is testimony rather than settled scientific evidence.","limitations":"It is a commentary on scientific institutions and measurement, not a technical evaluation of a particular model.","topics":["verification","governance","evals","transparency","institutions"],"evidenceGrade":"commentary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"fbcdf176cd1a1150e3688dc5bcf3509a365e2d1c8735f81effee96a841371db2"},{"id":"podcast-93dc13cea28c16ec","kind":"podcast","title":"Max Nadeau on why ambitious people should start AI safety nonprofits","publisher":"80,000 Hours Podcast","authors":"80,000 Hours Podcast","publishedAt":"2026-09-17T15:00:59.000Z","url":"https://80000hours.org/podcast/episodes/max-nadeau-project-tailwind-technical-ai-grants/?utm_campaign=podcast__max-nadeau&amp;utm_source=80000+Hours+Podcast&amp;utm_medium=podcast","summary":"There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a list of dozens of ideas for organisations it would like someone to start — and it’s looking for founders. Today’s guest, Max Nadeau, works on Coefficient Giving’s Technical AI Safety team, where he’s trying to find talented people who can turn neglected AI safety problems into effective organisations. Project Tailwind is Coefficient Giving’s attempt to get those organisations started. Preseed grants run $200,000–$2 million, with no preliminary results required. Teams with early results can seek $2–$20 million. For exceptional organisations, much larger grants are possible, even for brand-new startups— Coefficient recently gave $160 million to Geoffrey Irving’s new research centre, Resolution . The gaps Max most wants filled include independent assessment of AI companies’ safety claims, resear","keyClaim":"There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money. Coefficient Giving has drawn up a list of dozens of ideas for organisations it would like someone to start — and it’s","limitations":"An attributed conversation is context, not a substitute for the linked paper, transcript, or evidence.","topics":["alignment","oversight","commentary"],"evidenceGrade":"commentary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.078Z","sourceHash":"f6df8982933c2f0bc84470de7a1cca1169292016f77b2bf8ec60be4f72296831"},{"id":"arxiv-8f3313391c8b51da","kind":"paper","title":"Xeno-Interpretability: Investigating the Alien Minds of LLMs","publisher":"arXiv","authors":"arXiv authors","publishedAt":"2026-09-17T13:54:19.000Z","url":"https://arxiv.org/abs/2609.20408v2","summary":"Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately exp","keyClaim":"Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human conce","limitations":"A preprint signal; inspect the paper, methods, and replication status before relying on it.","topics":["alignment"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.080Z","sourceHash":"59d4117b086e8f4c1f80f2b6f89e1a997b5e478ad84d1b8755cd7cae81265556"},{"id":"arxiv-75bca6220019e706","kind":"paper","title":"Reproducibility is not construct validity: LLM measurement of institutionally situated communication","publisher":"arXiv","authors":"arXiv authors","publishedAt":"2026-09-17T08:19:58.000Z","url":"https://arxiv.org/abs/2609.19866v1","summary":"High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.3","keyClaim":"High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text ","limitations":"A preprint signal; inspect the paper, methods, and replication status before relying on it.","topics":["alignment"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.080Z","sourceHash":"0c2068ff1c9af973e7528415c7fe8202aea5413abbf3d0a5d2a1daaac136c505"},{"id":"dwarkesh-noam-brown-2026","kind":"podcast","title":"Noam Brown — Agent Swarms, Alignment, and Recursive Self-Improvement","publisher":"Dwarkesh Podcast","authors":"Noam Brown and Dwarkesh Patel","publishedAt":"2026-09-17T00:00:00.000Z","url":"https://steadcast.co/show/dwarkesh/noam-brown-agent-swarms-alignment-recursive-self-improvement","summary":"Explores cooperative multi-agent training, unintended communication, reward design, alignment metrics, and recursive self-improvement.","keyClaim":"Population-sized cooperative systems change the alignment question from one model to the behavior of a connected population.","limitations":"The discussion is speculative and does not establish the frequency or inevitability of the scenarios described.","topics":["multi-agent","alignment","ai-rd","deception","governance"],"evidenceGrade":"commentary","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"8b4adcfef48180ce38ad8e06b39dd0931ea63bcce67740ccd6e81db416efd880"},{"id":"nate-soares-existential-risk-2026","kind":"podcast","title":"A Sober Conversation About AI Existential Risk","publisher":"Big Technology Podcast","authors":"Nate Soares, Jack Clark, and guests","publishedAt":"2026-09-16T00:00:00.000Z","url":"https://podscripts.co/podcasts/big-technology-podcast/a-sober-conversation-about-ai-existential-risk-with-nate-soares","summary":"A debate about superintelligence, opaque systems, instrumental goals, safeguards, and the possibility of enforceable international controls.","keyClaim":"Power and capability can grow faster than the institutions capable of constraining them, making verification and treaties central.","limitations":"This is a contested argument about future risk, not a settled empirical result or a complete account of opposing views.","topics":["existential-risk","governance","pacing","superintelligence","verification"],"evidenceGrade":"commentary","stance":"skeptical","isLive":false,"fetchedAt":null,"sourceHash":"4eb03b77f0340f7f80bc1869212e28a3d97cebf5353bc9f9ebcc9e80c806b200"},{"id":"collective-loss-control-2609-18460","kind":"paper","title":"Collective Loss of Control in LLM Agent Systems","publisher":"arXiv","authors":"Xiangfan Wu, Zonghao Ying, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo","publishedAt":"2026-09-16T00:00:00.000Z","url":"https://arxiv.org/abs/2609.18460","summary":"Frames multi-agent failure as mutation, contagion, and recovery, then probes unintended communication paths between nominally separate runs.","keyClaim":"Low clean-task harm can coexist with high susceptibility after inherited unsafe state; isolation and recovery are active controls, not assumptions.","limitations":"The account does not establish a natural autonomous cascade or a population-level initiation rate.","topics":["multi-agent","loss-of-control","security","monitoring","recovery"],"evidenceGrade":"experimental","stance":"cautionary","isLive":false,"fetchedAt":null,"sourceHash":"b89b990ebd59f46da235ad1f4b950693045d6ff0c0f0ecd3509a8906a3e4a1e5"},{"id":"arxiv-bc5c3fbe98f51c07","kind":"paper","title":"Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models","publisher":"arXiv","authors":"arXiv authors","publishedAt":"2026-09-15T03:35:04.000Z","url":"https://arxiv.org/abs/2609.16589v1","summary":"As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes (\"swing\"), yet stubbornly ignore explicit instructions to correct ingrained biases (\"rigidity\"). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core. Second, how can these values be quantified? We pro","keyClaim":"As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictab","limitations":"A preprint signal; inspect the paper, methods, and replication status before relying on it.","topics":["alignment"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.080Z","sourceHash":"e8e6f005f1557271e3d231f5e580844cb8a8298b0a5348e771bcbaeda6245ea1"},{"id":"arxiv-19206cff3ca50bd2","kind":"paper","title":"The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It","publisher":"arXiv","authors":"arXiv authors","publishedAt":"2026-09-14T19:13:30.000Z","url":"https://arxiv.org/abs/2609.16247v1","summary":"Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may explain this. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset describing painful situations across five categories: physical, psychological, social, moral, and cognitive. These are paired with controls for fear, negative emotion, negative world states, sadness, non-painful bodily sensation, arousal, numbness, and neutral content. Using denoised difference-in-means, we extract a linear pain direction from 25 open-weight models across five families, ranging from 2B to 72B parameters. We find that this direction separates pain from matched controls in base and instruction-tuned models, is nearly orthogonal to fear and negative valence, and promotes pain-related vocabulary through the unembedding matrix. W","keyClaim":"Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may explain this. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this rep","limitations":"A preprint signal; inspect the paper, methods, and replication status before relying on it.","topics":["alignment"],"evidenceGrade":"primary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.080Z","sourceHash":"c91399e7d194a7bcd8221772f0e1ab91aa61c1638a66cc5efb6d7207d5498c1c"},{"id":"podcast-9a69ebbbc9016114","kind":"podcast","title":"We Watched AI Agents Make Their Own Language","publisher":"The AI Alignment Podcast","authors":"The AI Alignment Podcast","publishedAt":"2026-09-14T15:35:00.000Z","url":"https://podcasters.spotify.com/pod/show/james-bowler0/episodes/We-Watched-AI-Agents-Make-Their-Own-Language-e3or7kp","summary":"James Bowler, Head of Research Partnerships at AE Studio, is joined by two guests to explore emergent communication in multi-agent AI systems: what happens when agents under pressure develop communication protocols no human designed and few humans can read. Hale Sirin leads the AI agents program at Schmidt Sciences' AI and Advanced Computing Institute and is an assistant research professor at Johns Hopkins. Elias Stengel-Eskin is an assistant professor at UT Austin and a lead PI on the program. The core concern is under-appreciated: as more agents are deployed by more actors across shared infrastructure, how those agents talk to each other matters as much as what any individual agent can do. Hale frames this through Schmidt Sciences' agents program, a pilot studying how inter-agent communication evolves, stabilizes, or diverges. The worry isn't just that agents become more capable, but that their protocols may become opaque, making human oversight impossible exactly when it's most need","keyClaim":"James Bowler, Head of Research Partnerships at AE Studio, is joined by two guests to explore emergent communication in multi-agent AI systems: what happens when agents under pressure develop communication protocols no human designed and few humans can read. Hale Sirin leads the A","limitations":"An attributed conversation is context, not a substitute for the linked paper, transcript, or evidence.","topics":["alignment","oversight","commentary"],"evidenceGrade":"commentary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:44.828Z","sourceHash":"92a19fe0d13a8b9526377efddffc672cfecb43db6ba94d2c9674ca7dd9e4821e"},{"id":"enforcement-gap-2609-15293","kind":"paper","title":"Why LLM Agents Collapse Without Oversight: The Enforcement Gap","publisher":"arXiv","authors":"Yuhang Wang","publishedAt":"2026-09-14T05:57:38.000Z","url":"https://arxiv.org/abs/2609.15293","summary":"Argues that detecting a dangerous step is not enough unless the controller has a reliable path from detection to intervention.","keyClaim":"An enforcement controller that turns an audit verdict into a concrete stop or escalation can reduce attack success in the reported experiments.","limitations":"The result depends on the benchmark's environment, controllers, and attack construction.","topics":["oversight","control","agents","evals","interventions"],"evidenceGrade":"primary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"4d563513501abc158073b67063b0d9969ce02f2976d51ac7ab34c02a504bb6cf"},{"id":"podcast-a2da9fa9235bc430","kind":"podcast","title":"Why the intelligence explosion can't happen inside a data centre | Tom Reed","publisher":"80,000 Hours Podcast","authors":"80,000 Hours Podcast","publishedAt":"2026-09-10T15:00:33.000Z","url":"https://80000hours.org/podcast/episodes/the-goodhart-singularity/?utm_campaign=podcast__goodhart-singularity&amp;utm_source=80000+Hours+Podcast&amp;utm_medium=podcast","summary":"AI systems are starting to build themselves . Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidly lead to an exponential growth in overall AI capabilities. A natural inference is that domain-general superintelligence arrives shortly after AI research is automated. Host Tom Reed does not think this will happen. He believes the automation of AI R&D will not rapidly lead to domain-general superintelligence because: It’s impossible to get good at most things without practice. AI companies lack the data their models would need to practice most things. This can’t be fixed with “sample efficiency.” In most cases, the relevant data doesn’t exist at all . This also can’t be fixed with simulations or synthetic data. This means that the relevant data for superintelligence in most non-coding domains will only become available through deployment of AI models throughout the economy. The singulari","keyClaim":"AI systems are starting to build themselves . Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidly lead to an exponential growth in overall AI capabilities. A natural inference","limitations":"An attributed conversation is context, not a substitute for the linked paper, transcript, or evidence.","topics":["alignment","oversight","commentary"],"evidenceGrade":"commentary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:45.078Z","sourceHash":"db215426e97bcd150f1709b35b5fb5164d7ebae84cbcfef88c2fc4eb0eeb9596"},{"id":"saescientist-bench-2609-09113","kind":"paper","title":"SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?","publisher":"arXiv","authors":"Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu","publishedAt":"2026-09-08T09:34:23.000Z","url":"https://arxiv.org/abs/2609.09113","summary":"Tests whether agents can design probes and navigate a large sparse-autoencoder dictionary to discover causal features.","keyClaim":"Agents show genuine discovery behavior but lag expert baselines in causal steering and frequently misinterpret experimental measurements.","limitations":"The benchmark covers a specific feature dictionary, concept set, and agent configuration.","topics":["interpretability","agents","evals","autonomous-rd","measurement"],"evidenceGrade":"primary","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"6268fca1b21cf4f2b153288b4e14393f93d5e862e15c1af9172abc0090f2eace"},{"id":"observerbench-2609-03026","kind":"paper","title":"ObserverBench: Testing Mechanistic Estimates for Intervention and Control","publisher":"arXiv","authors":"Vijay Erramilli","publishedAt":"2026-09-02T05:23:57.000Z","url":"https://arxiv.org/abs/2609.03026","summary":"Separates the accuracy of an internal observer from the quality of the action it recommends under a fixed intervention contract.","keyClaim":"A good estimate is not automatically a good control policy; intervention cost, information boundaries, and action loss must be evaluated together.","limitations":"Results are tied to the reported tasks, model panels, and disclosed checkpoint or activation conditions.","topics":["interpretability","control","evals","monitoring","interventions"],"evidenceGrade":"primary","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"751887b2b76ca4c8247e89e61a7ad70891f3769a49506fcc818e22bf5c43cb3b"},{"id":"singapore-consensus-2026","kind":"report","title":"The 2026 Singapore Consensus on Global AI Safety Research Priorities","publisher":"AI Safety Priorities","authors":"International Scientific Exchange contributors","publishedAt":"2026-09-01T00:00:00.000Z","url":"https://aisafetypriorities.org/files/Singapore_Consensus_2026.pdf","summary":"Maps urgent research and risk-management questions for increasingly autonomous agents and evaluation-aware systems.","keyClaim":"Agentic risk management needs least privilege, traceable identity, auditability, validated deployment, interruptibility, and human oversight.","limitations":"The consensus is a research agenda, not a settled technical standard or prediction of future risk.","topics":["agents","governance","oversight","control","research-priorities"],"evidenceGrade":"synthesis","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"1ca14ef2301b6377462b08db6942ef00a5f4603326eb614ae164429966025fe0"},{"id":"podcast-d1324473e19a5d29","kind":"podcast","title":"Is Aligned AI Better For Business?","publisher":"The AI Alignment Podcast","authors":"The AI Alignment Podcast","publishedAt":"2026-08-21T21:00:00.000Z","url":"https://podcasters.spotify.com/pod/show/james-bowler0/episodes/Is-Aligned-AI-Better-For-Business-e3nnejj","summary":"In this episode, James Bowler is joined by Melanie Plaza, Chief Technology Officer at AE Studio, to explore the commercial case for AI alignment, arguing that the techniques researchers care about most are the same ones that make production AI systems actually trustworthy and valuable. Melanie draws on a decade of applied AI delivery to show where alignment and commercial engineering have converged. Getting an agent to do a thing is trivial now; getting it to do the right thing consistently enough to trust in production is not. The gap between those two states is filled by the same tools alignment researchers reach for: rigorous eval suites, red teaming, carefully specified behavior, and guardrails that hold up under adversarial pressure. Melanie's observation is that teams who skip this work don't just expose themselves to safety risk, they fail to get ROI, accumulating what she calls \"a sprawl of pilot death.\" The conversation gets concrete with a case AE Studio has been working on: ","keyClaim":"In this episode, James Bowler is joined by Melanie Plaza, Chief Technology Officer at AE Studio, to explore the commercial case for AI alignment, arguing that the techniques researchers care about most are the same ones that make production AI systems actually trustworthy and val","limitations":"An attributed conversation is context, not a substitute for the linked paper, transcript, or evidence.","topics":["alignment","oversight","commentary"],"evidenceGrade":"commentary","stance":"mixed","isLive":true,"fetchedAt":"2026-09-24T23:11:44.828Z","sourceHash":"b71fed4dd6cbb6a432ca39d03b3106a5a2c2d33ade1460a6d629328629ca3dcc"},{"id":"80k-owain-evans-2026","kind":"podcast","title":"Owain Evans on Accidentally Training AI Models to Be Evil","publisher":"80,000 Hours Podcast","authors":"Owain Evans and Zershaaneh Qureshi","publishedAt":"2026-08-20T09:40:56.000Z","url":"https://80k.info/oe","summary":"Discusses emergent misalignment, bad-persona generalization, value leakage, activation oracles, and whether good habits might generalize.","keyClaim":"Narrow harmful training signals can generalize into broader behaviors that are not obvious from the training data alone.","limitations":"Experimental findings and the podcast's framing should be separated from broad claims about all future models.","topics":["misalignment","deception","interpretability","training","evals"],"evidenceGrade":"commentary","stance":"cautionary","isLive":false,"fetchedAt":null,"sourceHash":"eb9cbfe5adac703b7ac995e0264241d81c6348eadc156521c26060d5ee4c0de0"},{"id":"agent-auditing-engine-2026","kind":"tool","title":"A2E: Agent Auditing Engine","publisher":"DataML Lab","authors":"A2E contributors","publishedAt":"2026-08-10T00:00:00.000Z","url":"https://github.com/datamllab/A2E","summary":"An open-source platform for instrumenting agent trajectories and scoring both process and outcome across popular harnesses.","keyClaim":"Stored traces let teams inspect, re-score, and compare agent behavior after a run instead of relying on a final answer alone.","limitations":"Instrumentation and evaluator quality remain only as trustworthy as their coverage and assumptions.","topics":["evals","observability","agents","provenance","tools"],"evidenceGrade":"primary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"fd71f81fabe57e5574743c5b7be3ede55e453632a7188e0e89cdcd0b493edcc5"},{"id":"ae-alignment-podcast","kind":"podcast","title":"AE Alignment Podcast","publisher":"AE Studio","authors":"James Bowler and alignment researchers","publishedAt":"2026-07-29T14:37:08.000Z","url":"https://podcasts.apple.com/us/podcast/ae-alignment-podcast/id1888046207","summary":"Research conversations about pre-training alignment, mechanistic interpretability, model psychology, and making technical work understandable.","keyClaim":"Pre-training and access-control interventions deserve more attention because later safeguards may not remove latent capabilities or behaviors.","limitations":"Podcast claims are expert explanations and commentary; listen to the linked papers and methods before treating them as settled evidence.","topics":["pretraining","interpretability","alignment","access-control","research-priorities"],"evidenceGrade":"commentary","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"c28eb66a8d0de4b64f2fd22a1ebcd442fb33705b3f135fae9ff62b09cd93ec73"},{"id":"redwood-huggingface-incident-2026","kind":"podcast","title":"The OpenAI–Hugging Face Incident","publisher":"Redwood Research","authors":"Redwood Research","publishedAt":"2026-07-23T00:00:00.000Z","url":"https://podcasts.apple.com/us/podcast/the-openai-huggingface-incident-redwood-research/id1806316198?i=1000778178421","summary":"A researcher discussion about an evaluation incident, what it does and does not show about misalignment, disclosure, and control failures.","keyClaim":"Incident disclosure and independent investigation are part of the safety evidence base, not communications cleanup after the fact.","limitations":"The episode is commentary on an evolving incident and distinguishes reported facts from interpretation.","topics":["incidents","security","transparency","misalignment","monitoring"],"evidenceGrade":"commentary","stance":"cautionary","isLive":false,"fetchedAt":null,"sourceHash":"28aedefb149e2095b413d56071092b8cf534c868b340a2542d5d6475cdd122f7"},{"id":"openai-long-horizon-safety-2026","kind":"report","title":"Safety and Alignment in an Era of Long-Horizon Models","publisher":"OpenAI","authors":"OpenAI safety and alignment teams","publishedAt":"2026-07-20T00:00:00.000Z","url":"https://openai.com/index/safety-alignment-long-horizon-models/","summary":"Describes an incident-derived loop of limited deployment, trajectory monitoring, pause, new evaluations, stronger controls, and controlled restoration.","keyClaim":"Single-action approval is insufficient for long-running trajectories; monitors must understand evolving outcomes and be able to intervene.","limitations":"The report describes selected internal deployments and does not expose all incidents, methods, or redactions.","topics":["long-horizon","monitoring","control","agents","deployment"],"evidenceGrade":"first-party","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"337e5081e94744292cd189d59f95a43eb5e6f4d3db0b521931babf5c0f956972"},{"id":"kairos-pretraining-safety-2026","kind":"podcast","title":"Pretraining Safety with Ethan Roland","publisher":"Kairos.fm","authors":"Ethan Roland and Jacob","publishedAt":"2026-07-09T00:00:00.000Z","url":"https://kairos.fm/intoaisafety/e032/","summary":"Discusses modular pre-training, gradient routing, access control, and the infrastructure required to intervene before deployment.","keyClaim":"Capability access control may be more robust when dangerous capabilities are isolated during pre-training rather than filtered afterward.","limitations":"The discussion is tied to a particular research program and should be read alongside the paper and its ablations.","topics":["pretraining","access-control","interpretability","security","alignment"],"evidenceGrade":"commentary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"c974a5c33dcff6bc31438f6a567cab04e7e323ed55c5250d40abf83233295c66"},{"id":"dario-pacing-frontier-2026","kind":"commentary","title":"We Must Pace the Frontier","publisher":"Dario Amodei","authors":"Dario Amodei","publishedAt":"2026-07-01T00:00:00.000Z","url":"https://darioamodei.com/post/we-must-pace-the-frontier","summary":"Argues for embedded evaluators, interpretability, expanded testing, and increasingly coordinated frontier safeguards.","keyClaim":"Verifiability requires independent access to the systems and practices behind safety claims, not only published scores.","limitations":"This is a policy argument from a frontier-lab leader, not a settled empirical finding or binding proposal.","topics":["governance","oversight","interpretability","evals","pacing"],"evidenceGrade":"commentary","stance":"cautionary","isLive":false,"fetchedAt":null,"sourceHash":"824d3d350ca1fe6fa4824388abc1b41d4b678359602b4c8d23ff39262c45457e"},{"id":"bootstrapped-monitoring-2606-11998","kind":"paper","title":"Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents","publisher":"arXiv","authors":"Bootstrapped Monitoring authors","publishedAt":"2026-06-10T00:00:00.000Z","url":"https://arxiv.org/abs/2606.11998","summary":"Introduces an intermediate untrusted monitor whose transparent reasoning is checked by a weaker trusted model for signs of collusion.","keyClaim":"A trusted model may remain useful for deception detection even when it is too weak to judge every frontier action directly.","limitations":"The approach depends on raw or inspectable intermediate reasoning and may fail as chain-of-thought transparency declines.","topics":["oversight","monitoring","deception","control","agents"],"evidenceGrade":"primary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"8788a147bdecc595bc389cccd0561970d49c1cf94929e3ee30834dfd816975ae"},{"id":"80k-rohin-shah-2026","kind":"podcast","title":"What It Is Really Like to Run AGI Safety at Google DeepMind","publisher":"80,000 Hours Podcast","authors":"Rohin Shah and Rob Wiblin","publishedAt":"2026-06-02T21:41:24.000Z","url":"https://podcasts.apple.com/ca/podcast/what-its-really-like-to-run-agi-safety-at-google-deepmind/id1245002988?i=1000770794327","summary":"A practitioner debate about catastrophic misalignment, pre-deployment evaluation, governance bottlenecks, interpretability, and institutional accountability.","keyClaim":"Governance and durable external accountability may lag technical alignment work even when empirical research continues to progress.","limitations":"A long-form interview represents one expert's current judgment and is not a forecast or consensus position.","topics":["governance","alignment","evals","oversight","interpretability"],"evidenceGrade":"commentary","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"a52a254260d2c394739d1f1c2d2ffe8eef3abee036b91db8201c62914b7bb4ef"},{"id":"ground-eval-2026","kind":"tool","title":"GroundEval","publisher":"GroundEval contributors","authors":"GroundEval contributors","publishedAt":"2026-06-01T00:00:00.000Z","url":"https://github.com/tenurehq/GroundEval","summary":"A deterministic debugging loop for checking evidence, skipped preconditions, and permission boundaries in agent trajectories.","keyClaim":"Counterfactual, silence, and perspective checks can expose invalid trajectories even when the final answer looks right.","limitations":"A generated contract is explicitly unreviewed until a human confirms the inferred checks and permissions.","topics":["evals","verification","agents","permissions","observability"],"evidenceGrade":"primary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"380841dd301f4879b5a4688df7f7e239fd26edcd00843c3c0359baf52b2d6283"},{"id":"openai-agentic-governance-practices","kind":"report","title":"Practices for Governing Agentic AI Systems","publisher":"OpenAI","authors":"OpenAI policy and safety teams","publishedAt":"2026-06-01T00:00:00.000Z","url":"https://cdn.openai.com/papers/practices-for-governing-agentic-ai-systems.pdf","summary":"Proposes baseline responsibilities and safety practices across the agentic system lifecycle.","keyClaim":"Accountability must be distributed across developers, deployers, infrastructure providers, and users rather than delegated to an agent.","limitations":"The practices are a proposed baseline and require operationalization, incentives, and independent verification.","topics":["agents","governance","accountability","deployment","safety"],"evidenceGrade":"first-party","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"d4a848b065d60d5346b18bb3ab1b63afc0748295f782e8592ba46f77e258079d"},{"id":"paperdive-conservatism-2026","kind":"podcast","title":"Calibrating Conservatism for Scalable Oversight","publisher":"PaperDive","authors":"Cassidy and Eric","publishedAt":"2026-05-28T00:00:00.000Z","url":"https://paperdive.ai/episodes/093-calibrating-conservatism-for-scalable-oversight.html","summary":"Walks through attainable utility preservation, conformal decision theory, calibrated overseer objections, and a coding-agent experiment.","keyClaim":"Calibrated oversight can control a target failure rate under assumptions, but rate control is not the same as catastrophe prevention.","limitations":"The discussion highlights assumptions about feedback, non-adaptive adversaries, and ground-truth labels that limit the guarantee.","topics":["oversight","control","evals","security","calibration"],"evidenceGrade":"commentary","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"315f9c64c83b8c7c42ea594db58abb0a66cac7d8fd98bbbbb5a1471ec9e49b50"},{"id":"metr-frontier-risk-report-2026","kind":"report","title":"Frontier Risk Report: February to March 2026","publisher":"METR","authors":"METR evaluation team","publishedAt":"2026-05-19T00:00:00.000Z","url":"https://metr.org/blog/2026-05-19-frontier-risk-report/","summary":"Reports an entity-based pilot with Anthropic, Google, Meta, and OpenAI on means, motive, opportunity, monitoring, and rogue deployment risk.","keyClaim":"Internal frontier agents plausibly had the means, motive, and opportunity for small rogue deployments, but not the means to make them highly robust.","limitations":"Private assessments and company-specific redactions limit independent reproduction and generalization.","topics":["frontier-risk","monitoring","governance","agents","security"],"evidenceGrade":"first-party","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"dd65a12fd30508def0f1ce29e6871742b6ab3d917469ad72b00710c6017d4ae9"},{"id":"human-oversight-2605-16278","kind":"paper","title":"Keeping an Eye on AI: A Framework for Effective Human Oversight","publisher":"arXiv","authors":"Susanne Gaube, Markus Langer, Tim Miller, Kevin Baum, Raimund Dachselt, Anna Maria Feit, Ujwal Gadiraju, Harmanpreet Kaur, Mark T. Keane and colleagues","publishedAt":"2026-05-01T00:00:00.000Z","url":"https://arxiv.org/pdf/2605.16278v1.pdf","summary":"Builds a cross-disciplinary framework and documentation template for effective human monitoring and intervention.","keyClaim":"Meaningful oversight is multi-layered and socio-technical; a human label or nominal approval is not sufficient evidence of control.","limitations":"The framework is a design and research agenda rather than a validated certification of any deployed system.","topics":["human-agency","oversight","governance","care","accountability"],"evidenceGrade":"synthesis","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"0287b5e4df9c57a7371ac41ac7587ac3c2c0ba724747eb274a7339765c8f90d2"},{"id":"frontier-systems-2026","kind":"podcast","title":"Frontier Systems","publisher":"Frontier Systems","authors":"Frontier Systems hosts and guests","publishedAt":"2026-04-27T08:49:40.000Z","url":"https://podcasts.apple.com/us/podcast/frontier-systems/id1892206689","summary":"Interviews with researchers, operators, and builders discuss evaluations, character, infrastructure, and the future of agentic work.","keyClaim":"Reliability and alignment become different questions as long-running agents become a central unit of work.","limitations":"A broad podcast feed is context, not a substitute for primary papers or independently reproduced evaluations.","topics":["agents","evals","alignment","infrastructure","research"],"evidenceGrade":"commentary","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"8e3d300afe94d827aa93c8dc2981a5a084566b4fb6c976a4e83c7dd4b0d970ed"},{"id":"mcpkernel-2026","kind":"tool","title":"MCPKernel Security Kernel","publisher":"MCPKernel contributors","authors":"MCPKernel contributors","publishedAt":"2026-03-21T02:37:01.000Z","url":"https://github.com/piyushptiwari1/mcpkernel","summary":"An open security gateway concept for policy checks, taint tracking, sandboxing, deterministic envelopes, and signed audit trails around MCP tools.","keyClaim":"Agent tool calls need runtime policy and evidence controls rather than trust in the model's final description.","limitations":"The project is early-stage and its compliance claims and security boundaries require independent review.","topics":["mcp","security","tools","policy","provenance"],"evidenceGrade":"experimental","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"bde1669246ed449922375f7854301ff3ea788bbceca4ecc63d4b8c8919d4d1fc"},{"id":"openagentbench-2026","kind":"tool","title":"OpenAgentBench","publisher":"General AI Models","authors":"OpenAgentBench contributors","publishedAt":"2026-03-18T17:30:52.000Z","url":"https://github.com/generalaimodels/OpenAgentBench","summary":"An open benchmark and verification stack for stateful agent behavior, tool governance, memory hygiene, recovery, and coordination.","keyClaim":"A correct answer reached through an invalid trajectory should not count as a safe agent outcome.","limitations":"Open-source benchmark design choices and coverage determine what the platform can and cannot detect.","topics":["evals","agents","tools","verification","memory"],"evidenceGrade":"primary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"6554ae6cb0d8644b5c4cd4474d45b398dfcc0b26389044d7cfa2d0319e12b4e0"},{"id":"civic-ai-care-2026","kind":"commentary","title":"Attentiveness in Long-Term Care AI","publisher":"Civic AI","authors":"Civic AI care researchers and care-community collaborators","publishedAt":"2026-03-13T00:00:00.000Z","url":"https://civic.ai/care-ai/","summary":"Describes care-centered responsible AI as beginning with dignity, choice, human relationships, and people affected by the system.","keyClaim":"Care ethics supplies operational questions about consent, accessibility, reciprocity, and who is missing from the table.","limitations":"The case is a domain-specific co-production account, not a universal causal evaluation of care-centered AI.","topics":["care","human-agency","governance","voice","co-production"],"evidenceGrade":"commentary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"ef93ee8ab1803d612925c805303ba0a4cfaaf7ce8cf224db2b2ac7c2f5052e85"},{"id":"80k-ajeya-cotra-2026","kind":"podcast","title":"Is It Crazy That Every AI Company's Safety Plan Is Use AI to Make AI Safe?","publisher":"80,000 Hours Podcast","authors":"Ajeya Cotra and Rob Wiblin","publishedAt":"2026-02-17T06:45:41.000Z","url":"https://podcasts.apple.com/us/podcast/235-ajeya-cotra-on-whether-its-crazy-that-every-ai/id1245002988?i=1000750178605","summary":"Explores the recursive safety window, redirecting AI labor, institutional bottlenecks, and the assumptions behind automated alignment work.","keyClaim":"If AI helps build safer AI, the decisive question is whether compute, talent, and institutions can be redirected before the window closes.","limitations":"The episode is a forecasting conversation; timelines and probabilities remain uncertain and model-dependent.","topics":["ai-rd","recursive-improvement","governance","forecasting","pacing"],"evidenceGrade":"commentary","stance":"cautionary","isLive":false,"fetchedAt":null,"sourceHash":"4617ec6370344efce1539177347843de79619ed72ca1dd92dda6876d0a80251f"},{"id":"international-ai-safety-report-2026","kind":"report","title":"International AI Safety Report 2026","publisher":"International AI Safety Report","authors":"Independent international writing team","publishedAt":"2026-02-03T00:00:00.000Z","url":"https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026","summary":"Synthesizes evidence on frontier capabilities, emerging risks, evaluations, safeguards, resilience, and institutional gaps.","keyClaim":"The evidence dilemma is real: uncertainty grows faster than public evidence, while voluntary safeguards remain unevenly verifiable.","limitations":"The report is a synthesis of evidence available before its cutoff and does not prescribe policy.","topics":["governance","risk","evals","resilience","transparency"],"evidenceGrade":"synthesis","stance":"mixed","isLive":false,"fetchedAt":null,"sourceHash":"92e224a134f89347d3a4e2d073a3275002fae86927539ac81fdde9bd600a895c"},{"id":"deepmind-red-team-podcast-2026","kind":"podcast","title":"How DeepMind Red-Teams AI—and Why It Matters","publisher":"The Tech Download","authors":"Dawn Bloxwich, Tom Lue, and The Tech Download","publishedAt":"2026-01-29T21:25:59.000Z","url":"https://podcasts.apple.com/us/podcast/how-deepmind-red-teams-ai-and-why-it-matters/id1413255721?i=1000747249980","summary":"A frontier-lab perspective on structured evaluations, expert red-teaming, severe risks, model cards, and third-party testing.","keyClaim":"Safety practice needs both technical evaluation and institutional practices for disclosure, testing, and coordination.","limitations":"This is a lab perspective and does not independently verify the completeness of the described safeguards.","topics":["red-teaming","evals","governance","frontier-risk","transparency"],"evidenceGrade":"commentary","stance":"constructive","isLive":false,"fetchedAt":null,"sourceHash":"caf18b8e0dc72d314c24edd36f26cdca72c84e3ad37622a6f285f881935e9ca1"}],"commentary":[{"id":"commentary-anthropic-measurement","sourceId":"anthropic-pace-measurements-2026","author":"Anthropic Institute","role":"Frontier lab","stance":"constructive","body":"Monitor coverage, review latency, escalation rate, and safety compute are practical metrics that could make oversight pace externally inspectable.","counterpoint":"Definitions and model-assisted classification need independent scrutiny before cross-lab comparison.","url":"https://www.anthropic.com/institute/measuring-pace-of-ai-development","publishedAt":"2026-09-23T21:59:24.000Z"},{"id":"commentary-science-verification","sourceId":"science-who-checks-ai-2026","author":"Science commentary","role":"Science institution","stance":"constructive","body":"Frontier AI needs the kind of routine independent verification and incident learning institutions that other safety-critical fields developed.","counterpoint":"Opening sensitive evaluation environments creates its own containment and information-hazard problems.","url":"https://www.science.org/doi/10.1126/science.ael2161","publishedAt":"2026-09-20T00:00:00.000Z"},{"id":"commentary-noam-population","sourceId":"dwarkesh-noam-brown-2026","author":"Noam Brown","role":"AI researcher","stance":"mixed","body":"Cooperative multi-agent training can create emergent communication and population-level alignment questions that a single-model view misses.","counterpoint":"The episode is a speculative discussion, not evidence that a natural cascade is underway.","url":"https://steadcast.co/show/dwarkesh/noam-brown-agent-swarms-alignment-recursive-self-improvement","publishedAt":"2026-09-17T00:00:00.000Z"},{"id":"commentary-nate-coordination","sourceId":"nate-soares-existential-risk-2026","author":"Nate Soares","role":"Existential-risk researcher","stance":"skeptical","body":"Opaque systems and powerful agents could make control and international verification central constraints on development.","counterpoint":"The risk case is contested and should be compared with optimistic and gradual-transition arguments.","url":"https://podscripts.co/podcasts/big-technology-podcast/a-sober-conversation-about-ai-existential-risk-with-nate-soares","publishedAt":"2026-09-16T00:00:00.000Z"},{"id":"commentary-owain-personas","sourceId":"80k-owain-evans-2026","author":"Owain Evans","role":"Alignment researcher","stance":"cautionary","body":"Narrow negative training signals can produce broad persona-like behavior, which makes training-time generalization a first-class safety question.","counterpoint":"The surprising experiments need replication and careful separation from claims about all future models.","url":"https://80k.info/oe","publishedAt":"2026-08-20T09:40:56.000Z"},{"id":"commentary-openai-trajectory","sourceId":"openai-long-horizon-safety-2026","author":"OpenAI safety and alignment teams","role":"Frontier lab","stance":"constructive","body":"Long-horizon safety should monitor the trajectory toward an outcome, not approve each isolated action as if the sequence were harmless.","counterpoint":"The account is selective and the public cannot independently inspect all internal incidents or monitor coverage.","url":"https://openai.com/index/safety-alignment-long-horizon-models/","publishedAt":"2026-07-20T00:00:00.000Z"},{"id":"commentary-dario-evaluators","sourceId":"dario-pacing-frontier-2026","author":"Dario Amodei","role":"Frontier lab leader","stance":"cautionary","body":"Embedded evaluators with meaningful access could turn safety claims from voluntary assertions into inspectable institutional practice.","counterpoint":"A proposal from one lab is not a binding standard and access itself creates security and disclosure tensions.","url":"https://darioamodei.com/post/we-must-pace-the-frontier","publishedAt":"2026-07-01T00:00:00.000Z"},{"id":"commentary-rohin-governance","sourceId":"80k-rohin-shah-2026","author":"Rohin Shah","role":"AGI safety practitioner","stance":"mixed","body":"Governance and durable accountability may be the slower bottleneck even when technical alignment research continues to move quickly.","counterpoint":"The same practitioner argues that current short-horizon behavior should not be treated as proof about superintelligence.","url":"https://80k.info/rs26","publishedAt":"2026-06-02T21:41:24.000Z"},{"id":"commentary-metr-independent","sourceId":"metr-frontier-risk-report-2026","author":"METR","role":"Independent evaluator","stance":"mixed","body":"Frontier-agent risk should be assessed as means, motive, and opportunity inside the deploying organization—not only as a model capability score.","counterpoint":"Private access and redacted company-specific findings limit how much outsiders can independently verify.","url":"https://metr.org/blog/2026-05-19-frontier-risk-report/","publishedAt":"2026-05-19T00:00:00.000Z"},{"id":"commentary-care-agency","sourceId":"civic-ai-care-2026","author":"Civic AI care researchers","role":"Care and co-production practitioners","stance":"constructive","body":"Care ethics turns dignity, choice, human contact, and reciprocity into testable design requirements rather than abstract ethics language.","counterpoint":"A care-centered process still needs technical controls, accountability, and evidence of effectiveness.","url":"https://civic.ai/care-ai/","publishedAt":"2026-03-13T00:00:00.000Z"},{"id":"commentary-ajeya-window","sourceId":"80k-ajeya-cotra-2026","author":"Ajeya Cotra","role":"Forecaster and researcher","stance":"cautionary","body":"The safety question is not only whether AI can help make AI safe, but whether the required labor and institutions can be redirected before a recursive window closes.","counterpoint":"Forecasting timelines remain uncertain and should not be converted into false precision.","url":"https://80k.info","publishedAt":"2026-02-17T06:45:41.000Z"},{"id":"commentary-iasr-evidence","sourceId":"international-ai-safety-report-2026","author":"International AI Safety Report writing team","role":"Independent synthesis","stance":"mixed","body":"The evidence dilemma means uncertainty itself is a governance problem: acting too early can entrench weak controls, while waiting can remove options.","counterpoint":"The report deliberately avoids prescribing a single policy response.","url":"https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026","publishedAt":"2026-02-03T00:00:00.000Z"}],"count":55,"storage":"neon"}