I build analytics teams — and the AI agents the whole company actually uses.
Director of analytics at consumer product companies. Experimentation is where my rigor comes from; lately I build agentic systems that run in production and get used daily.
What's D7 retention for users who saw the new onboarding flow?
querying Snowflake · checking 2 tables
New-onboarding cohort: 34.2% D7 retention vs. 31.8% control (n=18,412).
Source: retention_by_cohort · SnowflakeIllustrative example — not live data
Selected work
Agents I've shipped
Production systems built with my team at a consumer subscription company — not side projects. Each one is a real problem, how we solved it, and what changed because it existed.
A natural-language analytics agent the whole company could trust
“A confident wrong number is the only failure that matters.” Built to answer business questions in Slack — sourced answers over plausible guesses, live in daily use.
Read the write-up →An agent that analyzes every live A/B test, every day, with no human in the loop
“Automating the experiment check made our statistics invalid.” A daily health check that prompted a rebuild from fixed-sample testing to something that survives being run every day.
Can an AI agent run an ad-creative pipeline end to end?
“The agent isn't the bottleneck — permission is.” A feasibility assessment: four of six stages are easy engineering, the other two are gated by platform trust, not agent capability.
Read the write-up →How I work
How I think about analytics
I build and scale data analytics teams — heavy focus on product analytics and experimentation, and more recently on agentic tools that automate the repetitive parts of the job. I've built an analytics organization from scratch, led company-wide AI adoption, and shipped agentic systems that are now used in production every day.
An analytics team's success should be measured by how many decisions it actually changed — not dashboards shipped, not queries run. If a report isn't moving a decision, automate it, and put the team back on work that does.
Below: a live simulation of a two-arm A/B test — not real data, just the kind of read I'd bring to a real one.
Fig. 1 — A tongue-in-cheek A/B test: the effect of adding one experimentation-minded analyst (me) to a team that was already confident in its gut calls. Simulated, not real data — but the instinct behind it is real. Pre-registered. Not peer reviewed. Results may vary by org.
Toolkit
What I reach for
Building with AI
Shipping agents to production, not evaluating tools.
- Claude
- Claude Code
- MCP
- Cursor
Agent patterns
The part that actually differentiates a candidate, beyond the vendor name.
- Tool-use loops
- Retrieval & knowledge layers
- Eval harnesses & LLM-as-judge
- Prompt-as-data
- Human-in-the-loop trust models
- Guardrails & evidence contracts
Data stack
Built the modern stack from scratch, twice.
- Snowflake
- dbt
- Segment
- Amplitude
- Airflow
- Looker
Analysis & experimentation
Where the rigor lives, underneath the agents.
- SQL
- Python
- Sequential testing & confidence sequences
- Causal inference
- Incrementality & geo testing
Contact
Let's talk
Hiring, or want to debate a methodology? Email is the fastest way to reach me.