{"items":[{"id":"openalex-0-preclinical-evaluation-of-psma-targeted-","title":"Preclinical evaluation of PSMA-targeted ultrasound contrast agents in an orthotopic model of prostate cancer in rabbits.","summary":"A public research record surfaced through OpenAlex for stewardship review.","publishedAt":"2027-01-01T00:00:00.000Z","source":"OpenAlex","url":"https://doi.org/10.1016/j.bioactmat.2026.07.040","tags":["research signal","governance"],"isLive":true},{"id":"openalex-1-the-sociolinguistics-of-machine-identity","title":"The Sociolinguistics of Machine Identity: LLM Personality and Ideology Propagation","summary":"A public research record surfaced through OpenAlex for stewardship review.","publishedAt":"2026-12-31T00:00:00.000Z","source":"OpenAlex","url":"https://doi.org/10.17613/fdwhn-ejz93","tags":["research signal","governance"],"isLive":true},{"id":"openalex-2-research-on-attention-guidance-and-user-","title":"Research on Attention Guidance and User Autonomy in AI-Powered Immersive Environments","summary":"A public research record surfaced through OpenAlex for stewardship review.","publishedAt":"2026-12-31T00:00:00.000Z","source":"OpenAlex","url":"https://doi.org/10.17613/vpe2q-5qg42","tags":["research signal","governance"],"isLive":true},{"id":"openalex-3-the-development-of-japanese-medical-ethi","title":"The development of Japanese medical ethics education from 1950s and its lessons for China","summary":"A public research record surfaced through OpenAlex for stewardship review.","publishedAt":"2026-10-01T00:00:00.000Z","source":"OpenAlex","url":"https://openalex.org/","tags":["research signal","governance"],"isLive":true},{"id":"openalex-4-psychological-reliance-and-ethical-orien","title":"Psychological reliance and ethical orientation in entrepreneurial intentions among university students: an applied psychology approach to employment and innovation","summary":"A public research record surfaced through OpenAlex for stewardship review.","publishedAt":"2026-10-01T00:00:00.000Z","source":"OpenAlex","url":"https://openalex.org/","tags":["research signal","governance"],"isLive":true},{"id":"openalex-5-codesign-and-system-integration-as-inter","title":"Codesign and system integration as interacting demands in digital health for youth experiencing chronic pain.","summary":"A public research record surfaced through OpenAlex for stewardship review.","publishedAt":"2026-10-01T00:00:00.000Z","source":"OpenAlex","url":"https://doi.org/10.1097/pr9.0000000000001494","tags":["research signal","governance"],"isLive":true},{"id":"arxiv-0-an-open-pipeline-and-dashboard-for-syste","title":"An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice","summary":"Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with human graders comparable to human--human agreement ($κ= 0.78\\text{--}0.82$), and a blind audit finds that $83\\%$ of sampled transformations preserve the original harm. In a survey ($N = 21$), most participants report that scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings","publishedAt":"2026-09-23T16:13:36Z","source":"arXiv","url":"https://arxiv.org/abs/2609.28335v1","tags":["AI safety","research signal"],"isLive":true},{"id":"arxiv-1-safety-nudges-user-facing-interventions-","title":"Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness","summary":"Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on individual nudges. Participants found the tool useful, clear, and minimally disruptive, with nearly all users reporting an increased awareness of potential AI harms, though we found that this improved awareness alone did not necessarily lead to discernible behavioral changes. Our results suggest that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context, while highlighting the importance of relevance, calibration, and user control in nudge design for conversational AI safety.","publishedAt":"2026-09-22T16:12:08Z","source":"arXiv","url":"https://arxiv.org/abs/2609.26865v1","tags":["AI safety","research signal"],"isLive":true},{"id":"arxiv-2-behavior-is-not-enough-a-mechanism-based","title":"Behavior is Not Enough: A Mechanism-Based Evaluation of Social Norm Emergence in LLM Societies","summary":"Social norms cannot be identified from behavior alone: the same cooperative equilibrium may reflect shared expectations, strategic incentives, or simple imitation. Yet in multi-agent large language model systems, prior work largely treats behavioral convergence as evidence of norm emergence. In this work, we introduce an evaluation framework that measures agents' reported empirical and normative expectations in addition to behavioral convergence. Through controlled ablations, we test the effect of expectation elicitation and isolate two collective mechanisms central to theories of norm formation---social learning through interaction and social selection through network-based group formation. We further test the stability of these resulting dynamics under adversarial disruption across four LLM families. We find that eliciting expectations increases cooperative contributions, while social learning stabilizes behavior, and social selection reliably identifies cooperators but provides limited behavioral reinforcement. Following disruption, normative expectations and behavioral coordination recover differently. Together, these results show that similar cooperative outcomes can arise from different underlying social processes. By making expectations observable, our framework allows us to attribute each mechanism's contribution separately, offering designers of multi-agent systems a principled basis for selecting the social processes that sustain cooperation.","publishedAt":"2026-09-22T14:21:04Z","source":"arXiv","url":"https://arxiv.org/abs/2609.26481v1","tags":["AI safety","research signal"],"isLive":true},{"id":"arxiv-3-from-decorative-to-load-bearing-task-dif","title":"From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought","summary":"Chain-of-thought (CoT) monitoring is only meaningful if written reasoning causally constrains the answer. We introduce continuation-based causal testing, an ablation-patch intervention that perturbs one reasoning step, truncates the chain, and forces the model to continue from the corrupted prefix. It measures how load-bearing a CoT is for the final answer, a behavioral notion distinct from mechanistic faithfulness. Across Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness tracks model-relative task difficulty: on easy tasks models silently bypass their own reasoning; on hard tasks they follow corrupted steps and propagate errors. A matched 2x2 analysis shows task difficulty dominates perturbation type: error propagation rises 16x from GSM8K to BBH multistep arithmetic, and a variance partition over 28,584 continuations attributes 98.8% of explained deviance to task difficulty versus 0.8% to perturbation type. Reasoning-specific RL suppresses error propagation and compresses the gradient. A four-variant judge-sensitivity analysis and blind two-annotator study (n=500) show the error-propagation vs. non-propagation label is invariant to judge prompt, with perfect inter-annotator agreement (Cohen's kappa = 1.00). This gradient creates a structural problem for CoT-based oversight and AI safety monitoring: where the trace is easy to read it carries little signal, and where it matters errors propagate before a monitor can intervene. Linear probes on hidden states separate silent bypass, self-correction, and error propagation, but additive activation steering provides limited causal control, flipping only about 25% of error-propagation cases at best. Behavioral mode is readable but not reliably controllable.","publishedAt":"2026-09-21T20:06:06Z","source":"arXiv","url":"https://arxiv.org/abs/2609.25366v1","tags":["AI safety","research signal"],"isLive":true}],"fetchedAt":"2026-09-24T23:10:36.637Z","live":true,"sources":[{"name":"arXiv","live":true},{"name":"OpenAlex","live":true}]}