Research brief
Explore how untrusted inputs can influence an AI workflow with tool access in an isolated test environment. This is an experiment outline, not a published result.
Questions to explore
Where should instructions end and untrusted data begin? Can narrowly scoped tools and explicit approvals limit unintended actions? How can failures be reproduced reliably?
Planned artifacts
A test harness, threat assumptions, sanitized examples, observed outcomes, and limitations. Results will be added only after the experiments are run.