Every public company's story is hiding in plain sight โ in 10-Ks and 10-Qs that
almost nobody reads end-to-end. The good stuff is specific: a gross margin
inflecting, risk-factor language that wasn't there last quarter, a going-concern
sentence buried on page 60. I wanted that surfaced to me every morning without
me doing the reading.
So for the All Things Agentic Hackathon I built EDGAR Sentinel: an
autonomous agent on Google Cloud that wakes up at 6:30 every morning, scans SEC
EDGAR for new filings across a 30-company watchlist, reads them with a
two-model pipeline, remembers every prior filing, and emails me what changed โ
with a public dashboard for everything it knows.
- Demo video (4 min): https://youtu.be/jJYZ74b0gUs
- Live dashboard: https://edgar-sentinel-dashboard-69101307007.us-central1.run.app
- Code: https://github.com/brianmyers-ctrl/edgar-sentinel
The architecture in one breath
Cloud Scheduler โ Cloud Run Job โ an ADK orchestrator agent (Gemini 3.5)
whose tools are the pipeline stages โ SEC EDGAR (politely: declared User-Agent,
throttled) โ raw filings archived to Cloud Storage โ a section parser โ
Gemma (on its own Cloud Run service, via Ollama) writes triage notes โ
Gemini 3.5 on Vertex AI scores the filing against a five-pillar "Filing
Health Score" with schema-enforced JSON โ Firestore stores it โ a delta
engine compares against the company's prior filing and fires deterministic
alerts โ SendGrid emails the digest.
One design rule shaped everything: agentic control flow, deterministic
execution. The LLM decides what runs and writes the run report. Tested
Python decides what is true โ section slicing, score weighting, the alert
rule, state transitions. When a judge (or I) ask "why did this alert fire?",
the answer is a rule you can read, not a vibe.
Two models, two jobs
Gemini 3.5 Flash does the deep reading: five pillar scores with cited
rationale, extracted metrics, three decision-relevant highlights. Temperature
zero, pydantic schema enforced, composite recomputed in code so config โ not
the model's arithmetic โ is authoritative.
Gemma's job is deliberately smaller: read the risk-factors section and produce
a dozen terse triage bullets โ red flags, notable changes, tone โ that ride
along to Gemini as a second opinion. It runs scale-to-zero on CPU. My first
design had Gemma rewriting filing text; that was wrong in an instructive way
(below).
What actually broke (the fun part)
-
Gemma 4 is a thinking model. My triage calls returned empty strings
with
done_reason: lengthโ the model spent its whole output budget on hidden reasoning and never wrote the answer. Onethink: falselater, 33-second useful triage notes. - Ollama's default context window silently truncates. My "cleaned" MD&A came back 12ร smaller โ not cleaning, truncation. That failure convinced me to change Gemma's job from rewriting text to writing notes about text.
-
Inline-XBRL splits words across spans. Microsoft's 10-K renders "RISK
FACTORS" as
RIS K FACTORS, and repeats "Item 1A" as a page header through the whole section โ my "take the last heading match" heuristic found nothing. Fix: match headings with optional intra-word whitespace and take the match with the longest following body. Microsoft's risk section went from 0 to 80,000 characters. -
A use-after-free in Python. Creating the google-genai client inline
(
make_client().models.generate_content(...)) let the client get garbage-collected mid-request; its finalizer closed the HTTP pool:Cannot send a request, as the client has been closed.Cached singleton. -
Org policies bite. Our org restricts Vertex models
(
constraints/vertexai.allowedModels) and strips default service-account grants โ both showed up as cryptic 400s/403s. Both fixed with scoped, least-privilege IAM rather than hammer-sized grants.
Every one of these would have detonated during a live demo. Finding them on day
one and day four instead is most of what "production-minded" means.
Does it actually notice things?
The delta engine is the feature I'd defend in a knife fight. Because every
analysis persists in Firestore, each new filing is compared with the company's
prior one โ pillar by pillar โ and a deterministic rule (โฅ10-point move, band
change, or risk-pillar collapse) decides whether to alert. On the full
backfill it flagged, among others: Plug Power sliding Caution โ Distress
(cash down to $161.9M), Salesforce and Meta dropping out of Strong, and
Coinbase and AMC genuinely recovering. It also caught Apple's management going
cautious on component costs a quarter before it showed up anywhere else in the
filing โ a 10-point management-signal drop while the composite barely moved.
The numbers
30 companies ยท 58 filings analyzed ยท 8 live alerts ยท running unattended every
morning since August 14 ยท ~1 minute per filing ยท 20 unit tests ยท roughly a
dollar a day in cloud costs while idle-scaling to zero.
One last production note: even the demo video is Google AI โ narration by
Cloud Text-to-Speech (Chirp3-HD), soundtrack generated with Lyria 2 on
Vertex AI, and the screen captured while the real daily job ran live on
Cloud Run.
I created this piece of content for the purposes of entering the All Things
Agentic Hackathon. EDGAR Sentinel produces automated research summaries
derived from SEC filings โ not investment advice.













