01 / Writings
Read the long version.
A single shelf for independent papers and public writing from my work at Semgrep.
Kimi K3's Code Security Results Look Competitive — Until You Look at Precision
A benchmark review of Kimi K3 that looks past headline F1 scores at precision, recall, and performance on larger codebases.
Grounded or Gamed? We Audited Our Own Cyber Benchmark
An audit of whether cyber-benchmark scores reflect grounded reasoning or shortcuts, with a focus on stability and counterfactual behavior.
Evaluating GPT-5.6 Luna, Terra, and Sol Against GPT-5.5 for AI Code Security
A model comparison through the lens of AI-powered vulnerability detection, including precision, recall, F1, and cost per true positive.
We Have Mythos at Home: GLM 5.2 Beats Claude in Our Cyber Benchmarks
A benchmark of GLM-5.2 against frontier models for IDOR detection, separating bare-model performance from the effect of a security harness.
Towards Transparent Reasoning
A survey of explainable AI methods, including LIME, SHAP, integrated gradients, and emerging reasoning-chain approaches for exposing model behavior.
The Diffusion Revolution
A review of the evolution from GANs to diffusion-based systems, covering DALL-E, Stable Diffusion, Sora, and open-source generative video work.