
My name is Josh Christensen, a lifelong engineer living in Pittsburgh, PA with roots in Northern Utah. For the past several years I’ve been a software engineer at Amazon, building the production inference systems behind conversational AI, including work for Alexa and Amazon’s Nova line of generative models.
My most recent work surrounds productionalizing novel LLM inference and classification approaches coming out of Amazon applied science teams. I led a small team building a Triton inference backend that derives paralinguistic features like confidence, sentiment and language identification from the token-level outputs of a vLLM-served LLM. We took this from prototype to production: single-digit-millisecond p50 latencies, and a concurrency refactor that took sustained throughput from ~8 to ~80 requests per second on the same hardware. Along the way I’ve built retrieval-augmented personalization for speech recognition, LLM-agent tooling that cut our on-call investigation time from half an hour to minutes, and the AWS CDK infrastructure underneath all of it.
I got here by way of systems software. Before Amazon I spent four years at L3Harris writing safety-critical C++ for radio systems, and I hold a BS and an MS in Computer Science from the University of Utah, where I graduated cum laude through the Honors program as an undergraduate and led the controls team for Formula SAE Electric. Outside of work I enjoy a variety of personal engineering projects, including a self-hosted, multi-modal inference stack serving roughly 35 models across text, speech, image, and video on a single RTX 5090.
