by Haytham ElFadeel - [email protected] May 2, 2026

You can have a productive week, improve every metric on the dashboard, and still be moving in the wrong direction. The team ships, the model gets better, and the roadmap looks reasonable. Meanwhile, the problem worth solving may have changed.

That's what draws me to the advice of Jim Keller, Andy Grove, Richard Hamming, and others. They approached different problems, but kept coming back to similar questions: What actually matters? Why does the current approach work? What would make us change it?

This post brings together some of the ideas I find most useful from them, alongside my own notes on research and engineering.

1. Understand the mechanism

“They use a transformer.” “They switched to diffusion.” “They added RL.”

These statements tell us something about a system, but much less than we assume. They identify an architecture, a modeling approach, or a way of learning. They don't tell us what the system actually learned, or why it works.

John Jumper makes this point when discussing AlphaFold: we like putting systems into familiar boxes, then treating the highest-level label as the explanation for their success. Calling AlphaFold 3 a diffusion model is technically correct, but that description can obscure the important work happening throughout the system. The label becomes a shortcut past the thing we wanted to understand. (Interview)

His example from AlphaFold 2 makes this concrete. People associated its success with equivariance - the way its geometric computations respect rotations and translations. Yet removing the invariant point attention component produced a relatively small loss compared with the overall improvement over AlphaFold 1. It helped, but the ablation suggested there was more to the story. Other choices, including the representations and loss function, were much less visible in the story people repeated.

A component can be important, recognizable, and technically interesting without explaining most of the result.

Suppose an ML model improves after a team introduces a Transformer. Several things may have changed together: which objects can exchange information, how the input is represented, the training data, and the optimization procedure. Saying “Transformers understand interaction better” skips the interesting part. What information was previously unavailable? How does it propagate now? Which change was necessary for the improvement?

Taxonomy gives us a vocabulary. Mechanism gives us a way to reason.

Jim Keller makes a related distinction between following a recipe and understanding the process behind it. A recipe can produce excellent bread. Understanding the process helps when the ingredients, equipment, or desired result change (interview).

Engineering teams accumulate recipes too: add this loss, tune that parameter, put a special case here. Those recipes are valuable, until the problem changes and we aren't sure which parts still apply.

Before committing to an important design decision, I'd want to answer four questions: