Back to Etched questions
CodingSoftware Engineer

Event Loop with Thread Pool

Role: Software Engineer


Implement a multi-threaded event loop in C++ that dispatches tasks to a fixed pool of P worker threads.

Problem Statement

Design and implement a ThreadPool class that:

  • Accepts an arbitrary number of tasks (callables) via a submit() method
  • Distributes tasks across exactly P worker threads
  • Processes tasks concurrently with no data races
  • Shuts down cleanly — all submitted tasks complete before destruction

Requirements

cpp
class ThreadPool {
public:
    explicit ThreadPool(size_t num_threads);
    ~ThreadPool();

    // Submit a task to be executed by a worker thread.
    // Returns a std::future so the caller can await the result.
    template<typename F, typename... Args>
    auto submit(F&& f, Args&&... args)
        -> std::future<std::invoke_result_t<F, Args...>>;
};

Example Usage

cpp
ThreadPool pool(4);  // 4 worker threads

auto f1 = pool.submit([] { return compute_heavy_thing(); });
auto f2 = pool.submit([](int x) { return x * x; }, 42);

std::cout << f1.get() << "\n";  // blocks until result is ready
std::cout << f2.get() << "\n";  // 1764
// Pool destructor waits for all in-flight tasks to finish

Key Design Considerations

Task queue: Use a std::queue<std::function<void()>> protected by a std::mutex. Worker threads block on a std::condition_variable when the queue is empty.

Shutdown: Set a stop flag before joining threads. Workers check the flag after waking — they drain remaining tasks before exiting.

Future/promise bridge: Wrap the user's callable in a std::packaged_task so callers can .get() the return value or any thrown exception.

Thread safety: Every read/write to the shared queue must be inside a std::unique_lock. The condition variable must re-check the predicate after waking (spurious wakeups).

Skeleton

cpp
#include <thread>
#include <mutex>
#include <condition_variable>
#include <queue>
#include <functional>
#include <future>
#include <vector>

class ThreadPool {
public:
    explicit ThreadPool(size_t num_threads) {
        for (size_t i = 0; i < num_threads; ++i) {
            workers_.emplace_back([this] { worker_loop(); });
        }
    }

    ~ThreadPool() {
        {
            std::unique_lock<std::mutex> lock(mutex_);
            stop_ = true;
        }
        cv_.notify_all();
        for (auto& t : workers_) t.join();
    }

    template<typename F, typename... Args>
    auto submit(F&& f, Args&&... args)
        -> std::future<std::invoke_result_t<F, Args...>>
    {
        using R = std::invoke_result_t<F, Args...>;
        auto task = std::make_shared<std::packaged_task<R()>>(
            std::bind(std::forward<F>(f), std::forward<Args>(args)...)
        );
        std::future<R> future = task->get_future();
        {
            std::unique_lock<std::mutex> lock(mutex_);
            tasks_.emplace([task]{ (*task)(); });
        }
        cv_.notify_one();
        return future;
    }

private:
    void worker_loop() {
        while (true) {
            std::function<void()> task;
            {
                std::unique_lock<std::mutex> lock(mutex_);
                cv_.wait(lock, [this]{ return stop_ || !tasks_.empty(); });
                if (stop_ && tasks_.empty()) return;
                task = std::move(tasks_.front());
                tasks_.pop();
            }
            task();
        }
    }

    std::vector<std::thread> workers_;
    std::queue<std::function<void()>> tasks_;
    std::mutex mutex_;
    std::condition_variable cv_;
    bool stop_ = false;
};

Follow-ups

  • How would you add a priority queue so high-priority tasks run first?
  • How would you implement work stealing to improve CPU utilization when some threads finish early?
  • What happens if a task throws an exception? How does std::packaged_task propagate it to the caller?
  • How would you add a bounded queue to apply back-pressure when tasks are submitted faster than they're consumed?

Variant (independent report) — Periodic Task Event Loop with Task Lifecycle

A second reporter shared the actual code for the Core SWE intern first interview ("usually any job with 'Core' is cpp"). In this telling the event loop is single-threaded and OS-flavored rather than a thread pool: tasks are registered with a period, kept in a min-heap ordered by next fire time, and each task's callback returns a lifecycle decision — keep rescheduling for another period, or garbage-collect/drop. The reporter called both Etched questions hard ("ya after this i never touched cpp interviews again"); another member: "first looks like a nightmare tho ... i think i just prefer dsa to os". Code shared verbatim:

cpp
#include <iostream>
#include <chrono>
#include <functional>
#include <queue>
#include <vector>
#include <thread>

class EventLoop {
    using Clock = std::chrono::steady_clock;

    struct TaskResult {
        bool keep;      // reschedule for another period?
        bool garbage;   // drop and clean up?
    };

    struct Task {
        int period;                              // ms between fires
        Clock::time_point next;                  // next fire time
        std::function<TaskResult()> fn;          // returns lifecycle decision

        bool operator>(const Task& o) const { return next > o.next; }
    };

    // min-heap: soonest next-fire on top
    std::priority_queue<Task, std::vector<Task>, std::greater<>> pq;

public:
    void registerTask(int period, std::function<TaskResult()> fn) {
        pq.push(Task{period, Clock::now() + std::chrono::milliseconds(period), std::move(fn)});
    }

    // Runs until no tasks remain (or forever if tasks keep themselves alive).
    void run() {
        while (!pq.empty()) {
            Task t = pq.top();
            pq.pop();

            std::this_thread::sleep_until(t.next);   // wait for it to be due

            TaskResult r = t.fn();

            if (r.garbage) continue;                 // drop it
            if (r.keep) {
                t.next = Clock::now() + std::chrono::milliseconds(t.period);
                pq.push(std::move(t));               // reschedule
            }
        }
    }
};

int main() {
    EventLoop loop;
    int count = 0;

    loop.registerTask(200, [&]() -> EventLoop::TaskResult {
        std::cout << "task A fired\n";
        return {true, false};                        // keep running
    });

    loop.registerTask(500, [&]() mutable -> EventLoop::TaskResult {
        std::cout << "task B fired (" << ++count << ")\n";
        bool done = count >= 3;
        return {!done, done};                        // stop after 3, then GC
    });

    // Let it run for ~2 seconds via a self-terminating guard task.
    auto start = std::chrono::steady_clock::now();
    loop.registerTask(50, [&]() -> EventLoop::TaskResult {
        bool expired = std::chrono::steady_clock::now() - start > std::chrono::seconds(2);
        return {!expired, expired};
    });

    loop.run();
    return 0;
}

The same interview loop's second round was the quadtree spatial-index question — see quadtree-point-in-boxes-query.md.

Source: community report, August 2026