← Home
Odd Lots · June 5, 2026 · 31m

Inside Hudson River Trading's Blistering Token Burn

Hudson River Trading's head of AI discusses how the firm is deploying large language models at scale, seven months after our last conversation. The episode covers token costs, compute bottlenecks, memory pricing, employee spending on AI inference, and whether HRT might build its own chips to manage the exploding economics of running AI systems in production.

This summary was generated from show notes and public descriptions, not from a full transcript review. Details may contain inaccuracies.

Novel

Custom Chip Development as Cost Management
Hudson River Trading is considering building its own chips optimized for AI inference to manage token and memory costs that general-purpose GPUs cannot address.

Highlights

Memory Pricing, Not Compute, Is the Bottleneck
The constraint on AI deployment is no longer raw GPU compute but rather memory—the cost of maintaining context windows for inference.

Editorial

Token Burn at Scale Reveals Hidden AI Economics
As trading firms deploy AI systems across thousands of employees, the per-token costs of running LLMs become a material operating expense that rivals infrastructure spending.
AI-Induced Delirium
The episode touches on a phenomenon where AI systems produce outputs that are internally consistent but factually or logically wrong—a form of computational overconfidence.
AI Deployment in Finance Requires Rethinking Operations
The economics, risks, and infrastructure needs of AI in trading are fundamentally different from the lab environment where models are trained.

Misc

HRT employees are spending measurable amounts on tokens just running AI tools internally
Memory pricing is becoming a serious constraint—more so than raw compute
The firm is actively considering custom chip development to manage AI costs
Token burn at scale reveals the hidden economics of AI deployment in finance
Was this useful?