The InfoQ Podcast · InfoQ

Improving Valkey with Madelyn Olson

February 9, 2026·38 min·2 clips
Valkey commands execute in one microsecond, making latency meaningless — throughput is the only metric that matters.
Valky performance is not mainly about latency. Olson separates the tiny cost of a simple command inside Valky from the much bigger delay added by network hops. The engine can be very fast. A command may take about one microsecond, while a network hop can take hundreds of microseconds, and cross-AZ traffic can reach a millisecond. That changes the measurement game. When the host asks about performance, Olson points back to throughput instead of isolated latency numbers. Throughput shows where the engine actually runs out of room. Once it hits that limit, contention inside the engine turns into big latency spikes, so the useful question is how much work Valky can handle before that happens. The benchmarking setup is plain on purpose. Send heavy traffic at the engine, watch how much it handles, and compare each change against a known path. Valky has tooling for this. Olson names Valky benchmark as the current reference point, while also saying the team is looking at other options. The sharper work happens closer to the code. Because Valky includes C code, the team can run small microbenchmarks again and again, then compare each implementation step. That shaped the hash table rewrite. Each change could be checked against the old behavior, with attention to weird regressions instead of just a shiny top-line gain. Memory access is still a wall. Olson calls out time spent waiting on main memory as something the team should measure more often, including CPU counters for memory stalls and prefetching. The episode ends practically, with pointers to a higher-level blog post, deeper technical posts, and the Slack community for listeners who want the hash table details.

As heard by us

A grounded Valkey performance discussion about throughput, hash tables, benchmarking, and memory-access costs.

InfoQ turns Valkey performance into a concrete engineering discussion, centered on how an in-memory system is measured and improved when simple commands may take about a microsecond but real deployments add network hops, contention, and memory-access costs.

Read the full review in PlayNext →

Why you'd press play

If you care about what makes a cache engine fast, this one stays concrete.

Read the full recommendation in PlayNext →
Listen to the show on