CTO @ Wrodium · NLP/IR @ Berkeley Hearst Lab · GEO-16
CTO at Wrodium. NLP/IR researcher at Berkeley's Hearst Lab. Currently building the knowledge-freshness infrastructure that decides what AI engines cite — and writing autopsies for dead benchmarks.
Berkeley · SkyDeck · Hearst Lab
Paper · company · rack
CHASE @ COLM 2026
Our study of how content ecosystems drift when creators repeatedly optimize against LLM ranking signals — published at the Conference on Language Modeling.
Wrodium
GEO and knowledge-freshness infrastructure that decides which sources AI engines trust and cite. Co-founder & CTO.
Agents & iron
Harnesses for supervising long-running agents — verification, approval gates, readable traces — backed by a 4× L40S server for local inference and fine-tuning.
Six flagship builds. Every card links to something running in production.
The refill layer for AI agents — one API to buy AI credits, compute, gift cards, domains and eSIMs, with policy gates and human approval built in.

An open marketplace of tools for AI agents — describe the job and what correct looks like; it routes, pays, verifies every result, and only charges when the job is done. Now powering bidder-side federal bid intelligence.


Public AI compute for California. Clusters for community colleges, council standup, fund analysis.
Open ↗
A 16-factor framework for what AI answer engines actually cite — grounded in an empirical audit of real citations, not vibes.
Open →
A Three.js mystery box inside a live iPhone — tap it, watch it burst open, win your purchase back.
Open →The real dashboard with sample data — click around, nothing to install and no signup. Describe a content workflow in plain English; the Copilot builds the agent pipeline.
I study how agents fail across long, tool-using workflows — and how sandboxing, verification, readable execution traces, and human approval gates make them safer to delegate to.
At UC Berkeley's Hearst Lab, under Prof. Marti Hearst, I work on information retrieval and AI answer systems. My research examines how ranking systems decide what to surface, cite, and trust.
I contribute to the UC CalCompute Coalition as a research contributor focused on benchmarking and procurement standards. I also operate a four-GPU L40S server that supports local inference, fine-tuning, retrieval, and research workloads.
I build systems for supervising long-running agents: bounded execution environments, scoped tool access, reversible writes, verification checks, approval gates, and readable transcripts for debugging failures.
My core engineering principle is simple: delegate tasks that are inexpensive to verify. A difficult task with a reliable automated check is often safer to delegate than a simple task requiring complete human review.

I built and operate a 4U GPU server. The system supports local inference for Llama-class 70B models, LoRA and QLoRA fine-tuning, embeddings, reranking, vision workloads, and concurrent smaller-model serving.
As co-founder and CTO of Wrodium, I built marketer-facing workflows that could crawl a website, evaluate its visibility in AI answer engines, identify structural problems, generate changes, route them through human approval, publish through six CMS integrations, and monitor subsequent citation performance.
The central architectural challenge was translating inconsistent CMS content into a common representation without losing meaning. We moved toward maintaining a verified source of truth about each business and generating content from that structured knowledge.

The Wrodium team at Berkeley SkyDeck
"How Content Ecosystems Are Reshaped When Ranking Is the Only Target" — published at COLM 2026 with Qianwen Gao, Zichang Su, Yiwen Hou, and Leanid Palkhouski.
CHASE studies what happens when content creators repeatedly optimize against an LLM ranking signal. Across six domains, ranking became progressively less aligned with independently evaluated quality — even as documents became better optimized for the ranking system.
CHASE on GitHub →
"AI Answer Engine Citation Behavior: An Empirical Analysis of the GEO-16 Framework"
An empirical study of factors associated with source citations across ChatGPT, Gemini, Perplexity, and Claude, based on more than 18,000 observed citations.
Read GEO-16 on arXiv →
agents, tool use, harness engineering, sandboxed execution, human-in-the-loop workflows
LLM judges, trajectory evaluation, golden datasets, regression testing, attribution, verification systems
search, embeddings, reranking, RAG, citation analysis, information retrieval
GPU serving, local inference, fine-tuning, queues, observability, fault recovery
backend systems, workflow orchestration, CMS integrations, structured content systems

A VR flight-simulator hardware company. Built affordable 3D-printed cockpit components that undercut the industry by 10×.
Read the story →
Wrodium sponsored a room of QA leaders and founders at Intercom's San Francisco HQ, digging into the black-box problem in AI testing.

"Why Your Content Isn't Getting Picked by AI" — Leanid Palkhouski, co-founder @ Wrodium. San Francisco, December 15, 2025.
I like comparing notes with people making real systems — researchers, founders, and anyone fighting bureaucracy with code.