Research Interview: CUDA Kernels, Flash-Decode, and Transformer Internals
Role: Research
Frequency: Reported
Report of an impromptu in-person interview at Sail Research (an efficient-inference / AI-agents startup — "efficient inferencing for AI agents ... fully utilizing all parts of the GPU"). The candidate was told it was "just a stop by the office kinda thing" and was interviewed on the spot without having prepped inference topics.
Topics asked:
- CUDA kernels
- Explain flash-decode
- Photonics
- Reinforcement learning
- Transformer internals
- Reliability in (ML) systems
The candidate described the round as an intense technical grill across all of these areas. Expect deep questions on GPU inference internals even in informal settings.
Source: community report, June 2026