- Role
- Technical PM · team lead
- Company
- Sea × OpenAI Codex Hackathon Vietnam 2026
- Period
- Hackathon · 2026
- Stack
- OpenAI Responses API, Agents, Graph detection, Policy guard
Sea × OpenAI Codex Hackathon Vietnam 2026
The problem
Fraudsters use many accounts to place fake orders at a shop they work with, take the cash, and never repay. Each order looks normal on its own; the fraud only shows across accounts, shops, phones and addresses. The blunt fix, blocking anyone who shares a phone, hits families and renters and slows growth. That is the tension: fraud is the biggest threat to holding bad loans near 1%, and a heavy hand is its own cost.
- Monee loan book, Q2 2026 (+62% year on year)
- US$11.1B
- first-time borrowers in one quarter
- 5.3M
- of SPayLater loans used outside Shopee via QR
- >20%
- bad loans (90+ days overdue): the line to hold
- ~1%
Source: Guardline proposal deck (Sea × OpenAI Codex Hackathon Vietnam 2026), p.2, citing Sea's Q2 2026 earnings.
Team lead
My role
I led a team of three. The tech lead built the investigation backend, orchestration, evidence checks and action limits; the AI engineer built the agent runtime, rule writer, anomaly scoring and evals. I owned the product side: I turned fraud rules into specs and tests, wrote the fraud scenarios and user flows, and designed the three-minute demo. The decision I pushed hardest on: write down what the agent may not do before writing what it does.
AI flow
How it decides
Cheap signals first, the agent for judgement, and a code guard before any action. The agent never acts itself. Tap a step to follow its path.
The decision
Block shared phones, or investigate in tiers?
Cheap signals clear most orders with no AI. An agent works the hard cases and never acts itself. A code guard checks every proposed action against fixed limits.
Cost I accepted: Anything touching more than 10 accounts, or a novel case, goes to a person.
Source: Guardline proposal deck (Sea × OpenAI Codex Hackathon Vietnam 2026), p.2, p.4 and p.7.
From an order to a decision
System design- Learning loop
Tap a step to follow its path and read what happens there.
Policy guard
What it may, and may not, do
Trust came from the limits, not the model. Play the guard: propose an action, change how many accounts it touches, and see who is allowed to take it.
The agent's output is a verdict, never an action. Every claim has to cite a real record, and the code checks that before anything happens.
New tactics
How it adapts
When it meets a trick no rule covers
Three minutes
The demo
The demo opens on the hardest thing to see: a ring no single order reveals, then the decoy that must not be caught. Try both.
A ring no single order reveals
14 accounts sharing devices. Orders are held automatically; the credit freeze goes to the bank.
The decoy
A family sharing one phone and a busy rental house. The desk clears them; a blunt rule would not.
A new cash-out tactic, live
No rule fires. The graph catches it; a rule is drafted, and promoted; the next wave is caught.
Why you can trust it
An audit trail, a consistency flag between two analysts, and a kill switch that drops it to recommend-only.
It learns only from human decisions and confirmed outcomes, never from its own verdicts.
Results
- cases resolved without AITarget, synthetic data
- ~95%
- catch rate on a new tactic, after adaptingTarget, synthetic data
- 0 → 90%+
- genuine users blockedTarget, synthetic data
- 0
Source: Guardline proposal deck (Sea × OpenAI Codex Hackathon Vietnam 2026), p.8: "Targets, not results. All numbers are computed live on synthetic data."
