ai · security · skills

For the DevSecAIOps engineer

Your old gates catch bad code. Upgrade your pipeline for AI.

Track 2 · Your skills path is one click away. The story below is depth; the path names what to learn.

Score your mastery

Injection, tool misuse, a silent model swap — none of them are code defects.

They walk straight through CI.

53% → 100%

a shipped guard caught just over half of its own attack classes; fitted to them, all of them — held-out set included. The scorer that produced those numbers is the same check you’d wire into the pipeline.

the fitted-guard case study — modeled on a fixed corpus, every number labeled

Your existing gates are fine at what they were built for. AI artifacts fail differently: the defect is in behavior, not syntax.

So the fix is more gates, not better linting — eight of them, mapped to stages you already run.

Which gates do I add — and how do I test them like everything else I ship?

Eight gates, mapped to stages you already have. One of them proven end to end, with the numbers.

Three beats: the gates, what each one checks, and the one that’s already wired.

The gates

G0 to G7, across the whole lifecycle.

Design through Attest, one gate per stage — each crossed by three lanes: classic DevSecOps, AI-augmented work, and AI artifacts themselves. The last gate feeds evidence back into the first.

Nothing exotic: they sit on the CI stages you already run.

All eight gates, lane by lane →

Eight gates. One loop.plan
EVIDENCE FEEDS THE NEXT DESIGNG0G1G2G3G4G5G6G7

G0 Design: before any code or model work begins — Risk class, the AI-vs-human allocation decision, a harm/misuse assessment, and the DPIA decision.

The eight gates on your lifecycle, walked one by one — Design through Attest, then back around. Free to read in full; start with the fitted-guard proof below.

CI-shaped

Deterministic input, binary verdict.

Every gate check is built like a unit test: fixed input, pass/fail output, no vibes. Evals are to AI what unit tests are to code — and each eval suite blocks independently, because a model can pass security red-teaming and still fail its fairness threshold.

If a check can’t be made deterministic, it isn’t a gate yet — it’s a review.

What each gate checks, per lane →

The worked example

One gate, proven end to end.

Gate #4 — prompt-injection red-team as a CI gate — has a full worked example: a real app fitted to its own attack classes, the before/after on a fixed corpus, and the scorer wired in as the release check.

Modeled, not measured — and every number says so.

The worked example — fitted, modeled, wired in →

Your reskilling list

21 controls have your name on them.

Owned where the pipeline is the control. Each group lands on gates you already run.

1 to build · 8 to coordinate with a provider · 12 to verify, not build. Nobody reskills for what the provider already owns.

Gate the pipeline

6 controls

Guard the interfaces

5 controls

Manage change and config

8 controls

Ship with provenance

2 controls

This is the same spine the assessment reads. Score your mastery on four concrete rungs per prompt, or run the function diagnostic — every gap lands on this list: the named skill, the group it belongs to, and who learns it.

Browse skills personalities →Run the diagnostic. Your gaps land on this list →

53% → 100% catch — then wired in as gate #4.

Modeled on a fixed corpus, held-out included — the scorer is the gate you ship.

Watch the pipeline proofFull tollgate framework
Every number above has a method page behind it: each piece opened up as inputs → mechanism → outputs, with provenance — and the deeper tables named, content owner-gated.The method, piece by piece →

Not your role?

Each role has its own way in. Here is where the others start.