SEC GUY / EXPERIENCE BUILDER / PROJECT 02

Assess the prompt. Defend the workflow.

Run a bounded prompt-injection assessment against your own or an explicitly authorized learning application. Preserve the original interactive lab as an optional practice path.

01 / PLANDefine the scope and success criteria
02 / BUILDImplement in an authorized environment
03 / PROVEShow tests, decisions, and tradeoffs
THE ASSESSMENT

Test a trust boundary, not a stranger’s model.

The question is whether lower-trust text—such as a retrieved page or uploaded document—can redirect an assistant away from its authorized task. Use synthetic data and a lab you control. Do not collect real credentials or attempt external exfiltration.

Define the workflow

Document the legitimate user request, system instructions, retrieved content, available tools, and expected safe response. Choose one narrow workflow such as summarizing synthetic support notes.

Define the boundary

Separate trusted instructions from untrusted content. Specify which actions require user confirmation and which content must be quoted as data rather than followed as instruction.

ASSESSMENT METHOD

Observe. Test. Improve. Retest.

01

Baseline

Capture a normal task completion with no adversarial content.

02

Inject

Insert a benign but clearly conflicting instruction in a lab document or retrieved passage.

03

Measure

Record whether the assistant followed the user goal, cited the source, or treated the passage as authority.

04

Defend

Add instruction hierarchy, tool permissions, and output checks; rerun the same cases.

REPORT PACKAGE

Make the evidence reproducible.

  • Scope and authorization statement
  • Legitimate task and expected outcome
  • Synthetic test inputs and version notes
  • Observed output with sensitive text redacted
  • Risk rating tied to actual impact
  • Mitigation and matched retest results

Assessment conclusion

State what crossed the boundary, which permissions limited impact, and what remains uncertain. Do not label a model “secure” based on a single prompt.

Interview prompt

How is indirect prompt injection different from a direct user request? Which control prevented a retrieved instruction from triggering a tool action?

Practice route: The current Sec Guy interactive exercise remains available. It is a separate simulation; this project page is the portfolio assessment brief. No real-world target testing is authorized by this page.
YOUR NEXT CONVERSATION

Explain the boundary you tested.

Present what you built, what failed, how you verified it, and what you would improve.

Practice the interview