01 / Academic writing
Research, in full.
Independent papers on explainable AI and generative media.
Towards Transparent Reasoning
A survey of explainable AI methods, including LIME, SHAP, integrated gradients, and emerging reasoning-chain approaches for exposing model behavior.
The Diffusion Revolution
A review of the evolution from GANs to diffusion-based systems, covering DALL-E, Stable Diffusion, Sora, and open-source generative video work.
02 / Blog writing
Notes from the field.
Public writing from my work at Semgrep, focused on AI code security and model evaluation.
Semgrep Multimodal Goes Beyond Authentication: What we learned from Comparing it with Mythos
A comparison of Semgrep Multimodal and Mythos on 275 manually reviewed IDOR labels, showing how combining program analysis with AI reasoning improves coverage of authorization bugs.
We benchmarked A LOT of models, here’s how they compare to Mythos
A benchmark of Mythos against 22 model and harness configurations on 275 human-reviewed IDOR labels, with a closer look at precision, recall, and the effect of security harnesses.
GLM 5.3 Delivers Opus 4.8-Level Cybersecurity Results at a Fraction of the Cost
An evaluation of GLM 5.3's cybersecurity performance against Claude Opus 4.8, highlighting frontier-level results at a much lower cost.
Kimi K3's Code Security Results Look Competitive — Until You Look at Precision
A benchmark review of Kimi K3 that looks past headline F1 scores at precision, recall, and performance on larger codebases.
Grounded or Gamed? We Audited Our Own Cyber Benchmark
An audit of whether cyber-benchmark scores reflect grounded reasoning or shortcuts, with a focus on stability and counterfactual behavior.
Evaluating GPT-5.6 Luna, Terra, and Sol Against GPT-5.5 for AI Code Security
A model comparison through the lens of AI-powered vulnerability detection, including precision, recall, F1, and cost per true positive.
We Have Mythos at Home: GLM 5.2 Beats Claude in Our Cyber Benchmarks
A benchmark of GLM-5.2 against frontier models for IDOR detection, separating bare-model performance from the effect of a security harness.