Luis Cerqueira

I build AI that shows its work.

Engineer and co-founder of Kaval, where AI agents check their facts before they act. Before that, Microsoft and Tern. I design, build and ship the whole thing, from crawler to iPhone app.

  • Co-founder, Kaval
  • Previously Microsoft, Tern
  • 3,517 commits in the last 12 months

Selected work

Six things I built this year, with the numbers that prove they work.

Kaval

Tells AI agents when the facts they rely on have changed.

Kaval watches health insurers' policy sites, catches every change, extracts the rules that matter and points to the exact spot in the source PDF. An agent calls one endpoint and gets a verdict with a signed receipt.

9,904
policy documents captured from four insurers, every file hash-verified, for about $2 of model calls
32,305/32,305
evidence boxes landed on the right spot in a frozen test set
$2K→$120K
first customer, from a monthly pilot to an annual contract
CrawlingPDF extractionChange detectionSigned receiptsMCP
A public insurer policy PDF with Kaval's evidence box drawn around the billing code Q5171
A real insurer policy. Kaval's box sits on the exact code it extracted.

Tinta

A little face under your Mac's menu bar that hears its name, talks back and shows its work.

It listens for its name entirely on the Mac, speaks through a realtime voice model, remembers you across Mac, WhatsApp and SMS, and runs browser errands with a second model reviewing anything consequential. Every reply is logged with the tools it used, so you can ask it what it did.

32/32
wake words caught, with 0 false alarms in 256 tries
<1%
of one CPU core to keep listening
197
tests across the Mac app and its memory service
On-device speechRealtime voiceMemoryBrowser agents
Tinta
Tinta's face, rebuilt for this page. Move your cursor.

Sharper clicks

Made a small vision model better at finding buttons on a screen, with zero extra training.

A zoom-and-refine pass around GUI-G2-3B, tested on ScreenSpot-v2. I wrote up what failed too: reinforcement learning cost 15 points, distillation cost 1.5, and cleaner labels mattered more than more data.

+3.9 pts
accuracy on web interfaces
+2.2 pts
accuracy on icons, the hardest category
<1 s
per click in the live demo, down from 6 s
PyTorchvLLMGPU servingEvals
Four ScreenSpot-v2 samples: the base model's click marked with a red X, the refined click with a green check inside the target
Real benchmark samples. Red is the base model's miss, green is the refined click.

ShipIt

Owns a pull request until production is verified.

It follows each PR through failing CI, review comments and deployment, repairs what breaks, and confirms the service is healthy afterwards. A human still presses merge.

5 days
of building. It now runs in production.
295
tests passing in CI, plus live database and real-browser checks
$0.03
per automated repair in the live-model drill
TypeScriptAzureAgentsDeploy safety
PR #214 · fix null check in parserexample run
  1. ✓CI failed on a type errordetected
  2. ✓Repair committed, one linepatched
  3. ✓Review comment addressedresolved
  4. ✓Deployed to staging, then productionrolled out
  5. ✓Health checks green for 10 minutesverified
An illustration of what a run looks like.

Seen

Search anything you've seen on your Mac.

Press ⇧⌘Space and ask. Seen captures your screens, reads them on-device with Apple Vision and searches with SQLite full-text. An optional model can rerank results, but only text leaves the Mac, and it falls back to local search.

100%
of text recognition and search runs on your Mac
7 days
of history, then it forgets on purpose
MIT
licensed, with privacy and architecture docs
SwiftScreenCaptureKitVision OCRSQLite FTS5
Seen's search window showing results for 'Where was that note about the launch date?'
Seen finding a note from earlier in the day (demo content).

Rookly

A chess coach that turns your mistakes into drills.

Play a game, study the moments you got wrong, drill the pattern. Stockfish runs on the phone, mistakes become spaced-repetition cards, and there's no account and no server.

1,263
tests, the app suite runs in about 3 seconds
120K
curated puzzles, the same set for everyone
0
servers. Everything runs on the phone.
SwiftUIStockfish 17.1Spaced repetition
Rookly's home screen: a Board Vision lesson, a weekly streak and ten puzzles to start
Today's plan: one game, your review, ten puzzles.

Index

Everything else worth opening. Tap a row for the details.

    The log

    Every commit in my local repos since June 1, one dot per day. I build with coding agents, so the counts run high.

    Before this

    Intern to lead engineer to founder, in five years.

    1. 2024 — now

      Lumaa, Inc.Co-founder · San Francisco

      Raised $850K from Founders Inc, Neo and Thales Partners. Three products on one thesis: Lumaa matched companies to government RFPs, where I led two engineers; Nivel leveled construction bids and cut price error from 41.6% to 11.5%; and Kaval tracks insurer policies for AI agents.

    2. 2023

      MicrosoftSoftware engineering intern · Seattle

      Built full-stack features for three projects for government customers.

    3. 2022 — 2023

      TernIntern, then lead engineer

      Built it from the ground up: shipped the MVP in four weeks and grew it past 2,500 users.

    4. 2021

      MaoniSoftware engineering intern · London

      Built the client onboarding app from the ground up and cut upload time by about 75%.

    5. 2021 — 2024

      MinervaComputer science

      Left to co-found Lumaa.

    Say hi.

    Working on something hard?

    I'd like to hear about it. Tear off a tab and the email is yours.