Skip to content
Cortex SentinelHackathon build · 2nd place

Compliance that scales like software, not headcount.

Hackathon build: rules in plain English, tested before they ship, a human signs every decision.

Cortex Sentinel detection console on a laptop
Role
Product lead: research, specs, console UI and pitch
Company
Agentic AI Build Week 2026
Period
Hackathon · Agentic AI Build Week 2026
Stack
Go, React, PostgreSQL, Elasticsearch + yente

The problem

The insight

The bottleneck isn't detecting risk. It's changing the rules.

Cortex Sentinel pitch deck, p.3

GoTyme had just handed its banking customers a crypto wallet (6.5M users, the bank's own public figure). Every , withdrawal and cross-exchange transfer can raise an alert, so alerts grow with volume. A digital bank can't answer ten times the volume with ten times the analysts, and for a licensed crypto exchange one AML miss is a licence event, not a fine.

The tension sat in 2,000 alerts analysts had already resolved. Some alert types were cleared every single time: pure noise. One, high crypto-wallet risk, split roughly 50/50: a coin flip only a person should call. So the real question was never "can AI score this?" It was "which alerts are we allowed to automate, and can we defend that to a regulator on any date?"

Agentic AI Build Week 2026
2nd place
resolved alerts in the team's corpus (seed data)
2,000
GoTyme users in the brief (company's public figure)
6.5M

Source: Cortex Sentinel pitch deck, p.2, p.3 and p.16. The user figure is GoTyme's own.

Ownership

What I owned

Thao owned

This was a team build on top of the open-source Marble decision engine. I led the product side: I researched Vietnamese AML law and the , wrote a 14-file user-story set, and specced every workflow as a clickable mockup before any UI existed. Then I re-skinned the console to match those mockups and wrote the pitch. Teammates built most of the backend and AI plumbing.

Two decisions were mine. First, spec before build: each screen existed as a mockup with real Vietnamese AML cases (fictional people) so the team argued about the workflow, not the pixels. Second, label every claim: in the pitch each number is tagged live, prototyped or , and the deck says out loud which parts were the platform's. Filter the map below to see exactly what that meant.

Thao owned
  • Disposition copilot mockup: a case with a 0.86 confidence score, a recommended escalation and cited reason codes

    Disposition copilot: every reason cites its evidence

  • STR drafting mockup: a queue of suspicious-transaction report drafts ranked by deadline

    STR drafting: drafts ranked by legal deadline

Two of my seven workflow mockups, specced before the console was built. All names are fictional.

Rule authoring

How a rule is born

Today a new laundering pattern means an engineer, a ticket and a blind deploy. In the console an analyst describes the rule in plain English, an agent turns it into a typed rule and checks it against the data model, and a person saves it. Step through it yourself.

Platform

Before it ships, the new version runs beside the live one

Demo data
71%
19%

Hover a segment, or tap a name below, to read it.

A candidate rule version is scored on the same traffic as the live version, with no effect on customers. This is how its decisions split in the demo run.Source: Cortex Sentinel pitch deck, p.7 (shadow test run, candidate v2 vs live v1).

Scoring

How a rule decides

Fire the rules, watch the score move

Demo rules
45
ApproveReviewBlockDecline

Score 45 lands inReview

Each rule adds points; the total lands in a band. The preset is the deck's structuring example: three transfers just under the reporting threshold.Source: Cortex Sentinel pitch deck, p.19 (rule weights and bands) and p.8 (structuring example).

The triage gate

What the machine may close

This is the team's own layer. A model recommends, but a fixed, auditable gate decides, and it fails closed: if anything is uncertain, a person gets the alert. Try to get the coin-flip alert auto-cleared.

Team build

Three-lane triage on the backtest corpus

Backtest, sample data
  1. All alerts100%
  2. Auto-closed by the triage gate6.5%
  3. True positives among the auto-closed0%

    No real case escaped through auto-close.

Only alerts the gate could clear safely were auto-closed; the rest went to an analyst or were escalated.Source: Cortex Sentinel project story (team backtest on sample/seed data). Not production results.

Manual review load, baseline rules vs the gate

Backtest, sample data
100Baseline rules
81.6With the triage gate
-18.4%
At a held 90% recall the gate cut false positives by 20.3% and manual review by 18.4%, against baseline rules at 100% recall and 18.1% precision.Source: Cortex Sentinel project story (team backtest on sample/seed data). Indexed: baseline = 100.

The decision

Who is allowed to close an alert?

Chosen: Transparent rules with a written reason

A model recommends, but a fixed, auditable gate decides, and it fails closed: if anything is uncertain, a person gets the alert.

Cost I accepted: Only a small share of alerts is closed automatically: 6.5% in the backtest, on sample data.

Source: Cortex Sentinel pitch deck, p.15 (fails safe, guard chain) and p.19; project story (backtest on sample data).

What shipped

The console

  1. 01

    Detect

    Scenarios and rules run on every transaction, with versions you can replay for a regulator on any date.

    Scenario list in the detection console
  2. 02

    Write a rule in plain English

    The rule studio where the agent's draft lands for a person to review.

    Rule studio editor
  3. 03

    Watch the decisions

    Every decision is logged with its outcome and score.

    Detection analytics: decisions over time and score distribution
  4. 04

    Model your own data

    Transactions, accounts and Travel Rule messages are mapped once, so every rule and review reads the same source.

    Data model: transactions linked to accounts and Travel Rule messages

Agentic AI Build Week 2026

Demo day

On stage after placing second. What I'd do next: wire the scorer into the live decision path and add a four-eyes publish gate.

Results

place, Agentic AI Build Week 2026Public result
2nd
of alerts auto-closed by the triage gateBacktest, sample data
6.5%
true positives among the auto-closed alertsBacktest, sample data
0

What I would do next: wire the scorer into the live decision path and add a publish gate.

Source: Cortex Sentinel pitch deck, p.16; project story (team backtest on sample/seed data, not production results). Placement from the event result.

Next project