Event Loop with Thread Pool
Role: Software Engineer
Implement a multi-threaded event loop in C++ that dispatches tasks to a fixed pool of P worker threads.
Problem Statement
Design and implement a ThreadPool class that:
- Accepts an arbitrary number of tasks (callables) via a
submit()method - Distributes tasks across exactly
Pworker threads - Processes tasks concurrently with no data races
- Shuts down cleanly — all submitted tasks complete before destruction
Requirements
class ThreadPool {
public:
explicit ThreadPool(size_t num_threads);
~ThreadPool();
// Submit a task to be executed by a worker thread.
// Returns a std::future so the caller can await the result.
template<typename F, typename... Args>
auto submit(F&& f, Args&&... args)
-> std::future<std::invoke_result_t<F, Args...>>;
};Example Usage
ThreadPool pool(4); // 4 worker threads
auto f1 = pool.submit([] { return compute_heavy_thing(); });
auto f2 = pool.submit([](int x) { return x * x; }, 42);
std::cout << f1.get() << "\n"; // blocks until result is ready
std::cout << f2.get() << "\n"; // 1764
// Pool destructor waits for all in-flight tasks to finishKey Design Considerations
Task queue: Use a std::queue<std::function<void()>> protected by a std::mutex. Worker threads block on a std::condition_variable when the queue is empty.
Shutdown: Set a stop flag before joining threads. Workers check the flag after waking — they drain remaining tasks before exiting.
Future/promise bridge: Wrap the user's callable in a std::packaged_task so callers can .get() the return value or any thrown exception.
Thread safety: Every read/write to the shared queue must be inside a std::unique_lock. The condition variable must re-check the predicate after waking (spurious wakeups).
Skeleton
#include <thread>
#include <mutex>
#include <condition_variable>
#include <queue>
#include <functional>
#include <future>
#include <vector>
class ThreadPool {
public:
explicit ThreadPool(size_t num_threads) {
for (size_t i = 0; i < num_threads; ++i) {
workers_.emplace_back([this] { worker_loop(); });
}
}
~ThreadPool() {
{
std::unique_lock<std::mutex> lock(mutex_);
stop_ = true;
}
cv_.notify_all();
for (auto& t : workers_) t.join();
}
template<typename F, typename... Args>
auto submit(F&& f, Args&&... args)
-> std::future<std::invoke_result_t<F, Args...>>
{
using R = std::invoke_result_t<F, Args...>;
auto task = std::make_shared<std::packaged_task<R()>>(
std::bind(std::forward<F>(f), std::forward<Args>(args)...)
);
std::future<R> future = task->get_future();
{
std::unique_lock<std::mutex> lock(mutex_);
tasks_.emplace([task]{ (*task)(); });
}
cv_.notify_one();
return future;
}
private:
void worker_loop() {
while (true) {
std::function<void()> task;
{
std::unique_lock<std::mutex> lock(mutex_);
cv_.wait(lock, [this]{ return stop_ || !tasks_.empty(); });
if (stop_ && tasks_.empty()) return;
task = std::move(tasks_.front());
tasks_.pop();
}
task();
}
}
std::vector<std::thread> workers_;
std::queue<std::function<void()>> tasks_;
std::mutex mutex_;
std::condition_variable cv_;
bool stop_ = false;
};Follow-ups
- How would you add a priority queue so high-priority tasks run first?
- How would you implement work stealing to improve CPU utilization when some threads finish early?
- What happens if a task throws an exception? How does
std::packaged_taskpropagate it to the caller? - How would you add a bounded queue to apply back-pressure when tasks are submitted faster than they're consumed?
Variant (independent report) — Periodic Task Event Loop with Task Lifecycle
A second reporter shared the actual code for the Core SWE intern first interview ("usually any job with 'Core' is cpp"). In this telling the event loop is single-threaded and OS-flavored rather than a thread pool: tasks are registered with a period, kept in a min-heap ordered by next fire time, and each task's callback returns a lifecycle decision — keep rescheduling for another period, or garbage-collect/drop. The reporter called both Etched questions hard ("ya after this i never touched cpp interviews again"); another member: "first looks like a nightmare tho ... i think i just prefer dsa to os". Code shared verbatim:
#include <iostream>
#include <chrono>
#include <functional>
#include <queue>
#include <vector>
#include <thread>
class EventLoop {
using Clock = std::chrono::steady_clock;
struct TaskResult {
bool keep; // reschedule for another period?
bool garbage; // drop and clean up?
};
struct Task {
int period; // ms between fires
Clock::time_point next; // next fire time
std::function<TaskResult()> fn; // returns lifecycle decision
bool operator>(const Task& o) const { return next > o.next; }
};
// min-heap: soonest next-fire on top
std::priority_queue<Task, std::vector<Task>, std::greater<>> pq;
public:
void registerTask(int period, std::function<TaskResult()> fn) {
pq.push(Task{period, Clock::now() + std::chrono::milliseconds(period), std::move(fn)});
}
// Runs until no tasks remain (or forever if tasks keep themselves alive).
void run() {
while (!pq.empty()) {
Task t = pq.top();
pq.pop();
std::this_thread::sleep_until(t.next); // wait for it to be due
TaskResult r = t.fn();
if (r.garbage) continue; // drop it
if (r.keep) {
t.next = Clock::now() + std::chrono::milliseconds(t.period);
pq.push(std::move(t)); // reschedule
}
}
}
};
int main() {
EventLoop loop;
int count = 0;
loop.registerTask(200, [&]() -> EventLoop::TaskResult {
std::cout << "task A fired\n";
return {true, false}; // keep running
});
loop.registerTask(500, [&]() mutable -> EventLoop::TaskResult {
std::cout << "task B fired (" << ++count << ")\n";
bool done = count >= 3;
return {!done, done}; // stop after 3, then GC
});
// Let it run for ~2 seconds via a self-terminating guard task.
auto start = std::chrono::steady_clock::now();
loop.registerTask(50, [&]() -> EventLoop::TaskResult {
bool expired = std::chrono::steady_clock::now() - start > std::chrono::seconds(2);
return {!expired, expired};
});
loop.run();
return 0;
}The same interview loop's second round was the quadtree spatial-index question — see quadtree-point-in-boxes-query.md.
Source: community report, August 2026