Skip to main content

Technical Work

Open the code, rerun the pipeline, and check the numbers yourself.


What's here

These are projects you can inspect: each page links its source code, and most ship with a seed database so they run without an API key. Recent work comes first. I've kept the 2021 experiments up because they show where the statistics and software habits came from, including the ones that didn't work.

Recent systems

  • Strategy Game Market Study: 209,223 Steam reviews across 29 strategy games. RoBERTa-large beat VADER by 19.9 points of balanced accuracy on a test I set before reading the held-out reviews, and 24 findings held up when any one game was dropped.
  • AI-Verified Economic Analytics Pipeline: joins 10 FRED series with 1,547 Hacker News stories, computes every number in Python, lets a language model write only the prose, then checks each displayed value against SQLite.
  • NLP Text Analytics: the April baseline for the market study. VADER sentiment against Steam's recommendation flag across 2,382 reviews, with the 823 reviews too thin to model flagged and kept out.
  • Spreadsheet-to-Database Consolidator: turns messy spreadsheets into SQLite, validates every row, and quarantines rejects with their source file, row number, and a reason someone can act on.

Statistics and econometrics

Algorithms and scientific computing

Recent work without public code

  • Player Research and Engineering for a Strategy Game: a matched comparison of how players rate three battle turn systems, a study of which community posts gain traction, and the agent workflow behind an 81% faster pathfinding benchmark.
  • A Lead CRM and Call Ranker for a Wholesale Supplier: a bilingual lead CRM built on a city's open business data, with a random-forest ranker that raised closing rate per call about 50%.
  • A recommendation engine for a multi-source web data ingestion platform. The first model, a LightGBM model over sentence embeddings distilled from an LLM reviewer's labels, was validated once on a later held-out window at AUC 0.74 against 0.53 for a rule baseline. In replay, an MLP filter ahead of the LLM kept 60 of 74 top records with 719 LLM reviews instead of 1,707. More on the About page.

Related work

  • SicLib: the C++ scientific computing library behind the neural network demo, with linear algebra, numerical analysis, statistics, models, and a small neural network.
  • Articles: guides to pulling spreadsheet data out of Google Drive and choosing where it should go next.
  • Case Studies: client and studio work, with the method and result behind each one.