The journal / Applied AI Safety
Thoughtful ideas. Practical work.
Practical AI safety for the people building and shipping AI systems. Evaluations, permissions, prompt injection, and research put into practice.
6 articles
- ↗
A Retrieved Page Is Evidence, Not Permission
Trace an indirect prompt injection from source content to attempted action, then design boundaries that survive a mistaken model decision.
- ↗
Agent Memory Needs Data Boundaries, Not Just Better Retrieval
Design scoped memory access, preserve provenance, and account for derived notes, caches, deletion, and cross-agent sharing.
- ↗
An Incident Playbook for Tool-Using Agents
Contain risky actions, preserve useful evidence, reconcile uncertain effects, and restore service without replaying the incident.
- ↗
Design Evaluations That Can Support a Safety Claim
Choose observable outcomes, separate task success from harm, and read evaluation results without turning a narrow benchmark into a universal promise.
- ↗
Evaluate the Permission Boundary Before the Model
Build a permission matrix, test the executor, and distinguish blocked attempts from unauthorized effects.
- ↗
Make Human Approval a Decision, Not a Reflex
Design review around consequential actions, clear evidence, and approvals that cannot silently authorize a different operation.
No articles match. Try a different search or topic.