Builder · Founder · Researcher

Arlen
Kumar

CTO @ Wrodium · NLP/IR @ Berkeley Hearst Lab · GEO-16

I build products, platforms, GEO infrastructure, AI benchmarks, agent-readable surfaces, and the systems that fight bureaucracy and measure AI.

  • products.
  • platforms.
  • GEO infra.
  • benchmarks.
  • RAG systems.
  • agent layers.
  • web apps.
  • APIs.
  • backends.
  • pipelines.
  • dev tooling.
  • living articles.
  • cockpits.
  • foundations.

CTO at Wrodium. NLP/IR researcher at Berkeley's Hearst Lab. Currently building the knowledge-freshness infrastructure that decides what AI engines cite — and writing autopsies for dead benchmarks.

Berkeley · SkyDeck · Hearst Lab

UC BerkeleySkyDeck '24Prof. Marti HearstAir Quake (exit)SCET featured
UC BerkeleyCS · Data Science · Economics
Hearst LabIR + AI answers · Prof. Marti Hearst
Berkeley SkyDeck '24Backed · Most Innovative Tech
One exitAir Quake Simulations · acquired
● Now — Aug 2026

Three live
workstreams.

Paper · company · rack

CHASE @ COLM 2026

Our study of how content ecosystems drift when creators repeatedly optimize against LLM ranking signals — published at the Conference on Language Modeling.

PublishedCOLM 2026

Wrodium

GEO and knowledge-freshness infrastructure that decides which sources AI engines trust and cite. Co-founder & CTO.

$1.5M pre-seedSkyDeck

Agents & iron

Harnesses for supervising long-running agents — verification, approval gates, readable traces — backed by a 4× L40S server for local inference and fine-tuning.

192GB VRAMIn the rack
Things I build

Shipped, live,
and clickable.

Six flagship builds. Every card links to something running in production.

01 · Company

Wrodium

The infrastructure layer between brands and AI — share of voice, citations, and the content that wins them. $1.5M pre-seed, Berkeley SkyDeck.

Wrodium Pro product screenshot

All projects & research →

Live demo — try it right here

This is Wrodium, running live.

The real dashboard with sample data — click around, nothing to install and no signup. Describe a content workflow in plain English; the Copilot builds the agent pipeline.

Now

What I'm working on now.

AI systems and evaluation

I study how agents fail across long, tool-using workflows — and how sandboxing, verification, readable execution traces, and human approval gates make them safer to delegate to.

UC Berkeley research

At UC Berkeley's Hearst Lab, under Prof. Marti Hearst, I work on information retrieval and AI answer systems. My research examines how ranking systems decide what to surface, cite, and trust.

Public and independent compute

I contribute to the UC CalCompute Coalition as a research contributor focused on benchmarking and procurement standards. I also operate a four-GPU L40S server that supports local inference, fine-tuning, retrieval, and research workloads.

Selected engineering

Systems I build and run.

Agent evaluation and harnesses

I build systems for supervising long-running agents: bounded execution environments, scoped tool access, reversible writes, verification checks, approval gates, and readable transcripts for debugging failures.

My core engineering principle is simple: delegate tasks that are inexpensive to verify. A difficult task with a reliable automated check is often safer to delegate than a simple task requiring complete human review.

Agents run on tokens — 5,481,846,657 tokens used

Research compute infrastructure

I built and operate a 4U GPU server. The system supports local inference for Llama-class 70B models, LoRA and QLoRA fine-tuning, embeddings, reranking, vision workloads, and concurrent smaller-model serving.

  • 4× NVIDIA L40S GPUs — 192GB total VRAM
  • Dual AMD EPYC processors
  • 512GB–1TB ECC DDR5 memory
  • Enterprise NVMe and bulk storage
  • 25–100GbE networking

Wrodium's automation platform

As co-founder and CTO of Wrodium, I built marketer-facing workflows that could crawl a website, evaluate its visibility in AI answer engines, identify structural problems, generate changes, route them through human approval, publish through six CMS integrations, and monitor subsequent citation performance.

The central architectural challenge was translating inconsistent CMS content into a common representation without losing meaning. We moved toward maintaining a verified source of truth about each business and generating content from that structured knowledge.

The Wrodium team at Berkeley SkyDeck

The Wrodium team at Berkeley SkyDeck

Research

Published work.

CHASE

"How Content Ecosystems Are Reshaped When Ranking Is the Only Target" — published at COLM 2026 with Qianwen Gao, Zichang Su, Yiwen Hou, and Leanid Palkhouski.

CHASE studies what happens when content creators repeatedly optimize against an LLM ranking signal. Across six domains, ranking became progressively less aligned with independently evaluated quality — even as documents became better optimized for the ranking system.

CHASE on GitHub →
CHASE — COLM 2026 conference paper

GEO-16

"AI Answer Engine Citation Behavior: An Empirical Analysis of the GEO-16 Framework"

An empirical study of factors associated with source citations across ChatGPT, Gemini, Perplexity, and Claude, based on more than 18,000 observed citations.

Read GEO-16 on arXiv →
GEO-16 research framework
Technical focus

The stack I think in.

AI systems

agents, tool use, harness engineering, sandboxed execution, human-in-the-loop workflows

Evaluation

LLM judges, trajectory evaluation, golden datasets, regression testing, attribution, verification systems

Retrieval

search, embeddings, reranking, RAG, citation analysis, information retrieval

Infrastructure

GPU serving, local inference, fine-tuning, queues, observability, fault recovery

Product engineering

backend systems, workflow orchestration, CMS integrations, structured content systems

In the field

Where the work shows up.

Air Quake Simulations immersive flight-sim cockpit

Air Quake Simulations

Acquired

A VR flight-simulator hardware company. Built affordable 3D-printed cockpit components that undercut the industry by 10×.

Read the story
AI testing panel at Intercom San Francisco

AI testing night at Intercom SF

Sponsor

Wrodium sponsored a room of QA leaders and founders at Intercom's San Francisco HQ, digging into the black-box problem in AI testing.

GEO Conference 2025

GEO Conference 2025

Speaking

"Why Your Content Isn't Getting Picked by AI" — Leanid Palkhouski, co-founder @ Wrodium. San Francisco, December 15, 2025.

Let's build something

Building at the edge of AI, law, or evals?

I like comparing notes with people making real systems — researchers, founders, and anyone fighting bureaucracy with code.