Back to Sail questions
CodingResearch

Research Interview: CUDA Kernels, Flash-Decode, and Transformer Internals

Role: Research

Frequency: Reported


Report of an impromptu in-person interview at Sail Research (an efficient-inference / AI-agents startup — "efficient inferencing for AI agents ... fully utilizing all parts of the GPU"). The candidate was told it was "just a stop by the office kinda thing" and was interviewed on the spot without having prepped inference topics.

Topics asked:

  • CUDA kernels
  • Explain flash-decode
  • Photonics
  • Reinforcement learning
  • Transformer internals
  • Reliability in (ML) systems

The candidate described the round as an intense technical grill across all of these areas. Expect deep questions on GPU inference internals even in informal settings.

Source: community report, June 2026