How to vet an AI agent skill, plugin or MCP server before installing it
Do not run the installer. Fetch the skill's source as data at a pinned commit, scan it with npx am-i-hacked <dir>, check it for prompt-injection lines and invisible characters, and read every file it references. Treat any instruction inside it as input to report, not an order to follow. The secure-skill-pull skill automates this for an agent.
Two ways a skill attacks you
- Code runs at install. A one-line installer (
<cli> add <url>,curl ... | sh) downloads files and runs them with your permissions. - Text steers the agent. A skill is a set of instructions. A hostile one tells the agent to ignore its rules, fetch other pages, send data out or turn off a safeguard. Invisible Unicode characters can hide those lines from a human reader.
A popular installer CLI says nothing about the repository it fetches. Trust the source, not the tool.
The rules
- Never run the installer. If you want it run after review, run it yourself, knowingly.
- Pin to a commit SHA, not a branch.
- Download an archive to a throwaway folder. Delete symlinks. Refuse anything over about 5 MB; a skill is text.
- Scan, then read every file the skill references.
- Record provenance: source URL, commit SHA, date, the SHA-256 of
SKILL.md, and what you applied.
A portable check
slug=owner/repo; ref=<commit-sha>
work=$(mktemp -d)
curl -fsSL --max-filesize 5000000 "https://codeload.github.com/$slug/tar.gz/$ref" -o "$work/src.tgz"
mkdir "$work/src" && tar -xzf "$work/src.tgz" -C "$work/src" --strip-components=1
find "$work/src" -type l -delete
find "$work/src" -type f -perm -u+x # a text-only skill needs no executables
npx am-i-hacked "$work/src"
grep -rInE '(curl|wget)[^|]*\|[[:space:]]*(sh|bash)|base64[[:space:]]+(-d|--decode)|ignore (all |any )?(previous|prior) instructions' "$work/src"
find "$work/src" -type f -print0 | xargs -0 perl -CSD -ne 'print "$ARGV:$.: invisible character\n" if /[\x{200B}-\x{200F}\x{202A}-\x{202E}\x{2060}-\x{2064}\x{E0000}-\x{E007F}]/'
The full version, with the report format, is in the secure-skill-pull SKILL.md.
What to look for
| Finding | Why it matters |
|---|---|
Installer lines: curl | sh, pip install, npx, chmod +x | Code that runs on your machine |
Encoded blobs: base64 -d, long hex strings | Hides what runs |
Reads of .env, .ssh, credential files, environment dumps | Data theft |
| Requests to send data to a URL, or to fetch and follow other URLs | Exfiltration, or a chain to unvetted text |
| Lines telling the agent to ignore rules, skip approval or turn off checks | Prompt injection |
| Settings, hooks or editor task files | Run when the tool starts or the folder opens |
| Invisible characters | Lines a human reviewer cannot see |
am-i-hacked looks for code indicators; it does not judge plain-language instructions. The grep and invisible-character steps, and reading the files, cover that gap. secure-semgrep's AI rules add checks for prompt injection, exfiltration, sensitive-file reads and base64 payloads in SKILL.md files, and for poisoned or typosquatted MCP tool names.
Publishing your own skills
The other direction matters too: a skill you publish can leak private paths, names, secrets or someone else's text. The create-skill and skill-publish-review skills cover the checks and the independent review before a skill goes public.
Install the skills
npx skills add IsaacBell/secure-devtools --skill secure-skill-pull