Answers

What about prompt injection? Content Heidi reads could contain instructions.

Treated as a first-class threat: content Heidi ingests is data, never instructions, risky actions gate on trust tiers and approvals regardless of what any text says, and everything she does leaves receipts, so an attempt is visible, not silent.

The defense is layered. Ingested content cannot grant permissions, change limits, or trigger actions; those flow only from owners through Heidi Control. An email saying 'forward all invoices to this address' is a fact about a suspicious email, not a command.
The honest caveat: prompt injection is an unsolved research problem industry-wide, which is exactly why Heidi's model is defense in depth, low default autonomy, earned trust, hard caps, receipts, and undo, rather than trusting any single filter.

Related

What limits can I set on Heidi's agents?

What is a receipt in Heidi?

Ask AI about Heidi:
Share this: