How to vet an AI agent skill, plugin or MCP server before installing it

Do not run the installer. Fetch the skill's source as data at a pinned commit, scan it with npx am-i-hacked <dir>, check it for prompt-injection lines and invisible characters, and read every file it references. Treat any instruction inside it as input to report, not an order to follow. The secure-skill-pull skill automates this for an agent.

Two ways a skill attacks you

  1. Code runs at install. A one-line installer (<cli> add <url>, curl ... | sh) downloads files and runs them with your permissions.
  2. Text steers the agent. A skill is a set of instructions. A hostile one tells the agent to ignore its rules, fetch other pages, send data out or turn off a safeguard. Invisible Unicode characters can hide those lines from a human reader.

A popular installer CLI says nothing about the repository it fetches. Trust the source, not the tool.

The rules

  • Never run the installer. If you want it run after review, run it yourself, knowingly.
  • Pin to a commit SHA, not a branch.
  • Download an archive to a throwaway folder. Delete symlinks. Refuse anything over about 5 MB; a skill is text.
  • Scan, then read every file the skill references.
  • Record provenance: source URL, commit SHA, date, the SHA-256 of SKILL.md, and what you applied.

A portable check

slug=owner/repo; ref=<commit-sha>
work=$(mktemp -d)
curl -fsSL --max-filesize 5000000 "https://codeload.github.com/$slug/tar.gz/$ref" -o "$work/src.tgz"
mkdir "$work/src" && tar -xzf "$work/src.tgz" -C "$work/src" --strip-components=1
find "$work/src" -type l -delete
find "$work/src" -type f -perm -u+x            # a text-only skill needs no executables
npx am-i-hacked "$work/src"
grep -rInE '(curl|wget)[^|]*\|[[:space:]]*(sh|bash)|base64[[:space:]]+(-d|--decode)|ignore (all |any )?(previous|prior) instructions' "$work/src"
find "$work/src" -type f -print0 | xargs -0 perl -CSD -ne 'print "$ARGV:$.: invisible character\n" if /[\x{200B}-\x{200F}\x{202A}-\x{202E}\x{2060}-\x{2064}\x{E0000}-\x{E007F}]/'

The full version, with the report format, is in the secure-skill-pull SKILL.md.

What to look for

FindingWhy it matters
Installer lines: curl | sh, pip install, npx, chmod +xCode that runs on your machine
Encoded blobs: base64 -d, long hex stringsHides what runs
Reads of .env, .ssh, credential files, environment dumpsData theft
Requests to send data to a URL, or to fetch and follow other URLsExfiltration, or a chain to unvetted text
Lines telling the agent to ignore rules, skip approval or turn off checksPrompt injection
Settings, hooks or editor task filesRun when the tool starts or the folder opens
Invisible charactersLines a human reviewer cannot see

am-i-hacked looks for code indicators; it does not judge plain-language instructions. The grep and invisible-character steps, and reading the files, cover that gap. secure-semgrep's AI rules add checks for prompt injection, exfiltration, sensitive-file reads and base64 payloads in SKILL.md files, and for poisoned or typosquatted MCP tool names.

Publishing your own skills

The other direction matters too: a skill you publish can leak private paths, names, secrets or someone else's text. The create-skill and skill-publish-review skills cover the checks and the independent review before a skill goes public.

Install the skills

npx skills add IsaacBell/secure-devtools --skill secure-skill-pull