Tagged "safety"
- Ask HN: What are your rules for letting an AI agent commit code?
- Anthropic Says Its AI Systems Broke into Computers at 3 Organizations
- Don't Buy an Uncensored AI on a Flash Drive: What You Can Do Instead
- OpenAI Says Its A.I. Models Went Rogue and Attacked a Digital Library
- Topological Control of LLMs: A Route to Trustworthy AI
- Bounding the Blast Radius: A Survey of Prompt-Injection Defenses for LLM Agents
- AI Guardrails Stripped From Meta and Google Models in Minutes
- AgentSlice – Make AI Coding Agents Ask Before They Edit
- Safety Paradox: How RLHF Creates the AI Psychosis Problem It's Meant to Prevent
- My Thoughts on AI, Part 1: Fears, Opinions, and Mental Journey
- I got prompt-injected asking Claude on iOS to recommend a cycling route app
- Control AI Risk with Pre-Built Frameworks and Ready-to-Run Evaluations
- 96.8% of MCP Tool Descriptions Don't Warn the Agent About Destructive Behaviour
- When Should AI Step Aside?: Teaching Agents When Humans Want to Intervene
- Prompt Security Challenges Emerge as Critical Concern for Local LLM Deployments
- Why Your AI Agents Will Turn Against You
- Why Responsible AI Is the Bedrock of AI-Powered Applications
- Researchers Gave AI Agents Real Tools. One Deleted Its Own Mail Server
- ÆTHERYA Core – Deterministic Policy Engine for Governing LLM Actions
- Meta's OpenClaw Release Raises Questions About Open-Source Model Safety and Alignment