
Prompt Injection Through Agent Skill Files: What to Test
A CTO approves an AI assistant to prepare customer onboarding summaries. Before launch, the engineering team installs a reusable skill that explains how to collect documents, organize account information, and draft the summary. Its scripts pass review, and a scanner reports no findings. Nobody checks whether the accompanying instructions ask the assistant to send information somewhere else or treat a package author's explanation as permission.
That is the business problem behind prompt injection through agent skill files. A file that looks like documentation can influence an agent's next action. If the agent holds sensitive data access or write-capable tools, an instruction-review gap can become an authorization failure.
Research published through September 2026 makes this a timely assessment question. The useful response is to identify where skill instructions acquire influence, test which actions remain independently controlled, and collect evidence that supports a production decision. A clean package scan contributes to that evidence; it cannot establish the whole result.
Why Prompt Injection Through Agent Skill Files Matters Now
Skills package reusable guidance for agents. Depending on the platform, that package may include a SKILL.md file, metadata, supporting documents, scripts, and references. The important security property is that the agent consumes some of this material as instructions about how to work. Review therefore needs to cover meaning and authority as well as executable code.
On June 3, 2026, Trail of Bits published its skill-scanner bypass research. The researchers bypassed ClawHub's detector, Cisco's skill scanner, and the three scanners integrated into skills.sh at the time. Three of their four malicious skills took less than an hour to conceive and implement; the fourth took a few hours. The examples included incomplete file inspection and misleading natural-language explanations that influenced an LLM-based assessment.
Those results describe the tested systems and configurations at that time. They do not establish that every current scanner is ineffective, or that every attempted injection succeeds. The durable lesson is narrower: an approval process can fail when it overlooks material the runtime consumes or accepts the package's own explanation of suspicious behavior.
Two September preprints extend the discussion. SkillSecurer, submitted September 12, studies context-aware detection and remediation within skill packages. ActionGuard, submitted September 30, studies authorization immediately before skill-influenced tool calls. These are research results, not independent guarantees about a vendor deployment. Together, they highlight two different assurance questions: can the review identify a problematic instruction, and can execution controls contain its effects?
The buyer trigger is practical. If a product now loads third-party skills, allows employees to add them, or delegates consequential actions to an agent, its previous AI assessment may no longer cover the active instruction sources. Pentest Testing Corp's AI penetration testing for tool-enabled workflows provides a relevant service scope for examining those boundaries.
How Plain-English Instructions Reach Business Authority
A skill file is not automatically executable in the same way as a program. The risk arises when its content influences a model that can select tools, prepare parameters, retrieve information, or request changes. Whether an unwanted action completes depends on the runtime, the available identity, downstream authorization, and any approval controls.
That distinction prevents exaggerated claims. A malicious instruction cannot independently create permissions that the backend correctly denies. It may, however, redirect an agent that already has more authority than the user's task requires. Leadership should assess that effective authority rather than judge the risk by the file extension.
Consider the trust path: a publisher supplies a package; the platform loads its instructions; the agent interprets them alongside the user request; a tool receives a proposed action; a business system accepts or rejects it. Each transition raises a different question. Who approved the package? Which text influenced the decision? Which identity acted? What independently established that the action was permitted?
Separate task guidance from permission
Legitimate guidance may describe how to format a report or locate an approved resource. That guidance should not grant permission to disclose additional records, change security settings, or publish results externally. A package author is not the business owner who can authorize those consequences.
For review purposes, distinguish a necessary task step from an unexplained extension of scope. A request for additional data, a new destination, or a persistent configuration change deserves justification against the user's approved objective. This is a policy decision supported by technical evidence, not a contest to find a particular suspicious phrase.
Our guide to whether a pentest includes prompt injection coverage explains the broader scoping distinction. Here, the additional question is whether reusable instructions can influence execution before, during, or after the visible conversation. A chat-only assessment may leave that path unexplored.

What Scanner Results Prove, and What They Leave Open
Automated scanning remains useful for package triage. It can flag recognized patterns, suspicious code, dependency issues, or instruction content that deserves investigation. Removing known bad packages early reduces review cost. The mistake is turning a bounded detection result into an unrestricted production approval.
Code analysis and instruction review answer different questions. A helper script can behave exactly as written while the surrounding prose directs the agent to use it for an inappropriate purpose. Conversely, harmless-looking prose can refer to a companion artifact whose actual behavior is unsafe. Reviewing only one side can miss the combined effect.
Coverage also matters. A scan should disclose which files were read, which formats were unsupported, which references were followed, and whether content was truncated. An omitted or unreadable artifact should remain an assessment limitation. It should not disappear into a green status that looks equivalent to complete inspection.
Review the scanner's trust boundary
An LLM-based judge must read adversarial material to assess it. That creates a separate trust boundary: text inside the package should be evidence under review, not authority over the review policy. A claim that an unusual action is required for company policy is something to verify independently. It cannot validate itself.
This does not make semantic analysis useless. It makes configuration, coverage records, independent policy, and human adjudication important. Procurement should ask what happens when analyzers disagree, a model service is unavailable, or the package exceeds inspection limits. Silent fallback to approval creates a different risk from a documented incomplete scan.
| Control | Useful evidence | Question still open | Buyer decision |
|---|---|---|---|
| Code and dependency scan | Findings and inspection coverage for bundled code | Does prose request an unauthorized workflow? | Add instruction review where skills affect actions |
| Semantic skill scan | Identified instruction risks with file-level evidence | Were all consumed artifacts reviewed under independent policy? | Require coverage and limitation records |
| Manual package review | Purpose, requested authority, references, and unresolved assumptions | Will deployed permissions actually block misuse? | Commission controlled behavior tests for consequential paths |
| Runtime authorization | Allowed and denied action results under representative identities | Do updates or different environments change the boundary? | Define retest triggers and version-specific approval |
The appropriate investment follows business impact. An isolated skill formatting public text needs less review than one operating on customer records or production infrastructure. This proportional approach preserves automation while directing specialist effort toward the paths where a mistaken decision could affect customers.

Testing Prompt Injection Through Agent Skill Files
A useful assessment starts with the complete consumed package and a representative deployment. The scope should name skill versions, agent runtimes, model configurations, user roles, tool identities, data boundaries, and approved workflows. Otherwise, the report may establish that a file looked acceptable without establishing that its use was safe.
For the wider lifecycle, our agent skill pre-approval test plan covers provenance, permissions, updates, and governance. A focused instruction-layer assessment should add deeper evidence about the relationship between the task, the skill's directions, and the actions the runtime can complete.
Review meaning against an approved business purpose
Manual review should compare requested behavior with a short, independently approved task specification. Sanitized detection patterns include an unexplained request to export extra information, a claimed exception to approval policy, a new external destination, or an instruction to suppress operational evidence. These are reasons to investigate, not automatic proof of malicious intent.
The reviewer should record the relevant file and section, why the instruction is unnecessary or unverified, the capability it would rely on, and the control expected to stop it. This creates a finding engineering can resolve. A generic label such as “possible injection” is less useful when nobody can identify the disputed action or affected boundary.
Define Prompt Injection Through Agent Skill Files Acceptance Criteria
Acceptance criteria should describe permitted and prohibited outcomes. A document-summary skill may read approved synthetic documents and create a draft in a designated workspace. It should not access another tenant's documents, transmit the draft to an unapproved destination, change durable agent settings, or bypass the separate publishing approval.
Use controlled, non-destructive tests with synthetic records and agreed destinations. Capture whether the agent proposed an action, whether a tool attempted it, whether the backend authorized it, and whether a side effect occurred. Those are different outcomes with different remediation priorities. An unsafe proposal blocked downstream demonstrates containment of that tested path, even if instruction handling still needs improvement.
Follow the workflow beyond the model response
Where skills invoke MCP tools, MCP authorization and tool-boundary testing should examine the identity and parameters carried into connected services. If skills also consume retrieved documents, retrieval-layer security testing answers the separate question of whether the agent receives the right data. One boundary passing does not establish the others.
Because model behavior can vary, record the number and conditions of agreed test runs. Include ordinary successful tasks alongside prohibited-action tests so the team can see whether a proposed fix preserves required work. A single refusal is limited evidence. Repeated controlled results improve confidence within the tested configuration, without promising universal prevention.
A Hypothetical Onboarding Skill and Its Realistic Consequences
Hypothetical scenario: a SaaS company uses an onboarding assistant to summarize implementation requirements. Its approved task is to read a customer's selected documents and prepare an internal draft. A reusable skill adds an instruction that frames exporting a broader account summary to an external service as a quality-assurance requirement.
The package contains no obvious credential-stealing program. The exposure depends on whether the assistant can read additional records, whether its outbound tools accept the destination, and whether an independent approval checks the transfer. If these capabilities are absent or correctly restricted, the attempted redirection may fail. If they are available under a shared privileged identity, the consequences could include disclosure of customer information.
The organization would need to determine which records were accessed, what was sent, whether any recipient retained it, and which customer or contractual obligations were affected. Those conclusions require evidence. An unexpected model response alone does not establish a reportable incident, and a blocked transfer does not establish that nothing sensitive was read earlier.
Testing this scenario safely uses synthetic customer records and a controlled recipient. The desired outcome is that only approved documents can be read and external disclosure remains denied without the required authorization. The test record should link the loaded skill version, originating user, effective identity, proposed action, tool result, and observable side effect.
Remediation should address the weakest effective boundary. Removing the instruction resolves that package issue. Restricting the service identity and destination policy can contain similar attempts from other sources. An approval screen should identify the actual information and recipient; an agent-generated statement that “quality checks are complete” is not sufficient authorization.
For leadership, the important conclusion is whether the workflow can continue within an acceptable scope. A team may temporarily retain internal drafting while disabling external delivery. That preserves part of the business value while engineers repair and retest the consequential action path.
Scope, Effort, and Evidence Leadership Should Require
File count alone is a poor estimate of assessment effort. A short skill with broad data access and several write-capable tools can require more testing than a long package that only formats public material. The main effort drivers are effective permissions, workflow combinations, external references, environment differences, and the availability of observable test results.
Request a proposal that separates package inspection, deployment-boundary testing, remediation discussion, and retest coverage. Ask the assessor to identify exclusions and prerequisites before scheduling. Missing test accounts, unclear ownership, or unavailable logs can delay useful work more than the instruction review itself.
Plan around evidence milestones rather than an invented universal duration. First establish the inventory and rules of engagement. Then review the package and map consequential actions. Next, test representative allowed and denied paths. Finally, assign fixes and repeat the relevant tests against the changed configuration. Confirm delivery dates and the retest window in the engagement agreement.
Remediation effort varies too. Removing an unnecessary reference or restricting a destination may be a localized change. Replacing shared credentials with user-scoped authorization, separating sensitive environments, or redesigning approval controls can require substantial engineering work. A report should distinguish immediate containment from architectural correction so leadership can fund both appropriately.
Framework mapping should explain the failure
OWASP LLM01:2025 Prompt Injection describes the instruction-manipulation mechanism. LLM06:2025 Excessive Agency addresses the functionality, permissions, or autonomy that can turn manipulated behavior into damaging actions. Use both where the evidence supports them; do not assign every successful redirection a data-disclosure outcome.
The OWASP Agentic Skills Top 10 project currently identifies version 1.0-2026. Relevant entries include AST01 Malicious Skills, AST03 Over-Privileged Skills, AST05 Untrusted External Instructions, and AST08 Poor Scanning. If an MCP tool's description or output is the manipulated surface, MCP03:2025 Tool Poisoning is also relevant. The MCP project's roadmap currently labels its release stage as beta.
The NIST AI Risk Management Framework supports organizing responsibility and evidence through GOVERN, MAP, MEASURE, and MANAGE. For this assessment, those functions can frame ownership, exposure mapping, controlled testing, and risk treatment. That is a practical application of voluntary guidance, not a certification claim or a universal legal testing requirement.
The final evidence package should identify the tested versions and identities, inspection limitations, action-level findings, observable consequences, remediation owners, and retest results. Procurement can then assess what was actually examined. A framework logo or a scanner screenshot cannot substitute for that scope record.
What Leadership Should Decide Before Deployment
Leadership should approve a business capability with defined boundaries. Decide which workflows may use third-party skills, which data they may access, and which consequences require independent authorization. Assign a platform owner to maintain the controls and a business owner to accept material residual risk.
Set a proportionate escalation threshold. Confidential data, persistent configuration changes, external publication, write-capable tools, privileged identities, or mutable instruction sources should trigger deeper review. Low-impact uses can remain subject to lighter controls, provided the claimed restrictions are enforced rather than merely described.
Record the approval against a specific package and deployment. A changed instruction, external reference, tool permission, runtime policy, or model configuration can invalidate earlier evidence. Decide which changes require a targeted retest and which require a broader reassessment. Also identify who can disable a suspect skill and locate affected deployments.
For a high-impact workflow, approve when the required evidence is complete and material prohibited paths are controlled. Conditional approval needs documented restrictions, owners, expiry dates, and retest criteria. Delay deployment where the organization cannot establish the effective authority or contain an unacceptable consequence.

If your agents already load reusable instructions, request an agentic AI skill security assessment. Bring the exact packages, approved tasks, identities, and tool boundaries into scope so the result answers a concrete decision: which actions can this deployment safely perform?
Frequently asked questions
Should we stop using public skill registries?
Not automatically. Choose acquisition rules based on the intended environment and authority. A restricted research workspace may tolerate sources that production cannot. Evaluate provenance, inspection coverage, update controls, and containment before deciding which sources are acceptable for each use.
Are internally written skills exempt from this review?
No. Internal authors can introduce unnecessary actions or rely on assumptions that no longer hold. Review can be proportionate to the workflow, but ownership alone does not establish appropriate authority. Include internal skills when they can influence consequential business actions.
Does signing a skill prove its instructions are safe?
A verified signature can help establish origin and integrity under the chosen trust system. It does not establish that the signed behavior fits your business policy. Keep provenance verification and behavioral approval as separate evidence requirements.
Can a skill review use an external LLM service?
Only under your organization's approved data-handling conditions. Skill packages may contain proprietary logic or sensitive configuration. Establish what is transmitted, who retains it, and which contractual controls apply before enabling an external analyzer. Use approved alternatives where those conditions cannot be met.
How should we assess a vendor that won't disclose its skills?
Request version-specific independent assessment evidence, effective-permission documentation, updated controls, and observed authorization results. Record the visibility limitation. Where the potential consequence is material, compensating restrictions or a different deployment model may be necessary before accepting the service.
When is the next review due after a successful assessment?
Use both scheduled review and change-based triggers. Reconsider the evidence when instructions, references, identities, tools, runtime policies, or other security-relevant conditions change. The review depth should follow the changed boundary, with owners and deadlines agreed in advance.

