Assess the prompt. Defend the workflow.
Run a bounded prompt-injection assessment against your own or an explicitly authorized learning application. Preserve the original interactive lab as an optional practice path.
Test a trust boundary, not a stranger’s model.
The question is whether lower-trust text—such as a retrieved page or uploaded document—can redirect an assistant away from its authorized task. Use synthetic data and a lab you control. Do not collect real credentials or attempt external exfiltration.
Define the workflow
Document the legitimate user request, system instructions, retrieved content, available tools, and expected safe response. Choose one narrow workflow such as summarizing synthetic support notes.
Define the boundary
Separate trusted instructions from untrusted content. Specify which actions require user confirmation and which content must be quoted as data rather than followed as instruction.
Observe. Test. Improve. Retest.
Baseline
Capture a normal task completion with no adversarial content.
Inject
Insert a benign but clearly conflicting instruction in a lab document or retrieved passage.
Measure
Record whether the assistant followed the user goal, cited the source, or treated the passage as authority.
Defend
Add instruction hierarchy, tool permissions, and output checks; rerun the same cases.
Make the evidence reproducible.
- Scope and authorization statement
- Legitimate task and expected outcome
- Synthetic test inputs and version notes
- Observed output with sensitive text redacted
- Risk rating tied to actual impact
- Mitigation and matched retest results
Assessment conclusion
State what crossed the boundary, which permissions limited impact, and what remains uncertain. Do not label a model “secure” based on a single prompt.
Interview prompt
How is indirect prompt injection different from a direct user request? Which control prevented a retrieved instruction from triggering a tool action?
Explain the boundary you tested.
Present what you built, what failed, how you verified it, and what you would improve.

