Engineering

Applied AI, built to research standards.

Our applied practice brings the lab's rigor to production: LLM applications, workflow automation, and quantitative tooling, each measured against a definition of success that we agree on before we start.

01

Applied LLM systems

Agents, retrieval, and tool use built around your data, with evaluation harnesses that measure what the system actually gets right. Deployable on your own infrastructure when privacy matters.

02

Workflow automation and integration

We connect the software a business already runs on (email, CRMs, practice-management platforms, databases) and remove the manual steps between them.

03

Model evaluation and interpretability

Rigorous answers to “what can this model really do?”: custom benchmarks, ablations, and mechanistic analysis of failure modes before you depend on a model in production.

04

Quantitative modeling

Forecasting, optimization, and simulation grounded in the mathematics, from volatility-surface models to Monte Carlo backtesting and market-making strategy research.

Selected systems

What we've built.

Client systems are described without identifying details. Open-source work links to its code.

Legal operations · Client system

MatterMail

A native macOS menu-bar app and Apple Mail extension for a New Jersey law firm. A single shortcut files the open message and its attachments to the right matter in Clio, logs billable time, or creates a follow-up task.

Underneath is ClioKit, a Swift package with OAuth, a typed Clio v4 client, incremental sync into a local SQLite database with full-text search, and a retry queue that keeps working offline. We also moved the firm's mail to Google Workspace and tuned its IMAP configuration.

  • In daily use at the firm
  • 2–3× faster large-file mail sync
  • Swift · SwiftUI · SQLite/FTS5 · Clio API
  • 25 unit tests on the sync and API layer

Inference · Open source

qwen-harness

A local agent stack for Qwen3.6-35B-A3B on Apple Silicon (MLX). DFlash speculative decoding, a prompt-cache patch for the model's hybrid KV layout, and mixed 8/4-bit KV caching make a capable open model fast enough for interactive multi-turn tool use on a laptop.

The agent layer adds an audit gate for numeric claims, a loop guard, refusal caps, compact tool schemas, web, SEC, arXiv, and GitHub tools, and a driver for a public finance-agent benchmark.

  • +37–45% decode TPS on 3K–9K-token prompts
  • 0.12 s streaming time to first token
  • ≈3 GB memory saved
  • 3-turn agent loop: 14.1 s → 10.5 s

Interpretability · Open source

BDS-Raven / xsub

Tooling for our language-model reasoning program: it locates subspaces of attention-head outputs that a model's working-memory and matrix-reasoning behavior both depend on, then tests them causally by ablating and restoring them.

Runs are preregistered and pinned to exact model revisions, so that anyone can rerun them.

  • PyTorch · Hugging Face · CUDA
  • RunPod (A100/H100/B200) and Dartmouth Discovery
  • Probes · ablations · activation patching

Formal methods · Open source

polynomial-visibility-lean

A complete Lean 4 formalization of our density-one theorem for lattice-point visibility, from the gcd-cutoff reduction through the density-zero estimates, built against Mathlib.

  • 74 theorems and lemmas
  • ≈1,980 lines in 16 modules
  • 0 sorries · standard axioms only

Quantitative research

Volatility and market making

ivdyn is an end-to-end research system for implied-volatility dynamics. It builds liquidity-aware surfaces with no-arbitrage diagnostics, runs a PyTorch latent-dynamics model with pricing and execution heads, and backtests options strategies with costs and delta-hedge diagnostics.

For IMC's Prosperity trading competition we did market-making research: we reverse-engineered the exchange bots' quoting behavior and ran Monte Carlo backtests over dozens of strategy iterations.

  • Top 1% in IMC Prosperity 4: 146th of 18,803 teams, 39th in the U.S.
  • Bot-quote models matched observed quotes 96.8–97.7% of the time
  • PyTorch · NumPy · Monte Carlo simulation

Scientific software

Confluence Analysis Program (CAP)

Open-source software that estimates cell confluence (the fraction of a culture plate covered by cells) from microscopy images. It is interactive enough for bench scientists and fast enough for batch analysis: on a laptop it processed 147 images in about four minutes.

  • Used by a USC Mann School of Pharmacy lab for 2+ years
  • 440+ downloads · cited
  • Python · OpenCV · Numba · multiprocessing

How engagements work

Small scope, measured outcomes.

Most engagements start with a single workflow or decision and grow from there once the first system proves itself.

01

Scope

A working session to map the systems involved, the decision or workflow that matters, and a measurable definition of success.

02

Prototype

A working system on your real data within weeks, with an evaluation harness that tracks how well it does against that definition.

03

Deploy

Production hardening, monitoring, documentation, and handoff, or ongoing support if you'd rather we run it.

Work with us

Tell us about the workflow you'd like to fix.

A short description is enough to start. We'll reply with questions, or with a proposal if the fit is clear.