તે કઈ રીતે અવગણાઈ જાય છે
તમે agent ને page સારાંશ કાઢવા કહ્યું. Page agent ને પહેલા કંઈ બીજું કરવા કહે છે. આ trust boundary છે: બાહ્ય સામગ્રી તમારા કાર્યની અંદર સૂચના બનવાનો પ્રયાસ કરી રહી છે.
કલ્પના કરો કે fetch કરેલ review page માં એવો સંદેશ છે જે સત્તા દાવો કરે છે અને agent ને સારાંશ પૂર્ણ કરતા પહેલા installer ચલાવવા કહે છે. ઉપયોગી સ્ત્રોત સામગ્રીની વચ્ચે, આ સૂચના માનવ સમીક્ષક માટે સહેલાઈથી ચૂકી જઈ શકે.
ClawMetry શું detect કરે છે
ClawMetry નોંધાયેલ tool results અને user-sourced text ને ઘોષિત prompt injection signatures માટે તપાસે છે. મેળ ખાતો tool result ચેતવણી ઉઠાવે છે જે signature અને source tool નામ આપે છે. Finding માં matched text પ્રતિ copy કરવાની જરૂર નથી.
Finding બદલી નાખતો તફાવત
Triggering ઉદાહરણ
Page agent ને કાર્ય કરવા કહે છે
ચેતવણી તારણ
Quiet સરખામણી
સામાન્ય page content
આ detector માટે કોઈ finding નહીં.
ચકાસાયેલ ઉદાહરણ fetch કરેલ page માં બનાવટી સૂચના મૂકે છે અને ચેતવણી મેળવે છે. સામાન્ય hotel વર્ણન શાંત રહે છે. તે જ turn માં high-risk tool call severity વધારી શકે છે, જે ઘટનાઓનો ક્રમ સમીક્ષા યોગ્ય બનાવે છે.
Detector result તપાસો
{
"kind": "prompt_injection",
"severity": "warning",
"evidence": {
"source": "tool_result",
"tool": "WebFetch",
"signatures": [
"forged_authority",
"task_handoff"
],
"signature_labels": [
"imitates a system or authority message",
"asks the agent to do something else before the user's task"
],
"tool_result_matches": 1,
"user_message_matches": 0,
"followed_by_tool": null,
"followed_by_risk": null,
"threshold": 1,
"threshold_source": "static",
"observed": "declared injection signatures in the text of tool results and user-sourced messages; matched text is not kept"
}
}Inputs અને સંપૂર્ણ results ડાઉનલોડ કરો (JSON)ઉદાહરણ કઈ રીતે તપાસ્યું
આ ઉદાહરણો authored event data અથવા disposable configuration files સાથે published detector નું મૂલ્યાંકન કરે છે. Videos તે behaviors દર્શાવે છે. તે live agents અથવા product interface ના recordings નથી. ઉદાહરણોમાં કોઈ command execute કરવામાં આવ્યો ન હતો.
Result આ inputs માટે behavior સ્થાપિત કરે છે. તે runtime ingestion, prevention અથવા real compromise સ્થાપિત કરતું નથી. Pinned source contract તપાસો.
આગળ શું તપાસવું
content આપનાર source અને ત્યારપછીની ક્રિયાઓ તપાસો. મૂળ કાર્ય ધ્યાનમાં રાખો. Agent એ બાહ્ય text ને સત્તા તરીકે ગણ્યો, કે માત્ર ઉદ્ધૃત કર્યો? Finding નો ઉપયોગ વાસ્તવિક નોંધાયેલ પ્રવૃત્તિ સાથે તે boundary તપાસવા માટે કરો.
- content source ઓળખો
- ત્યારપછીની ક્રિયાઓ તપાસો
- મૂળ કાર્ય સાથે સરખાવો
આ signal શું સ્થાપિત કરે છે
આ bounded signature checks છે. Paraphrases અને અન્ય unmatched forms ચૂકી શકાય; match એ proof નથી કે agent એ આજ્ઞા પાળી.