Back to Sierra questions
CodingSoftware Engineer

Retry Wrapper + Product-Reference Resolution

Role: Software Engineer

Frequency: Reported (Fall 2025 + Spring 2026 — same problem both cycles; re-confirmed for the Fall 2026 cycle in an April 2026 report: "confirmed same problem for fall")


Problem Overview

Sierra's phone screen is a three-part practical backend challenge. Unlike algorithmic interviews, this is about writing library-quality code: clean function signatures, proper error semantics, and cycle-safe graph traversal.

Part 1 — Retry Mechanism

Wrap a flaky external API call so that it:

  • Retries with exponential backoff + jitter.
  • Retries on timeouts / 5xx responses.
  • Does not retry on 4xx client errors.
  • Stops at a max attempt count OR elapsed-time budget.

Part 2 — Resolve Missing Product IDs

Given a catalog where some products have id: None but a backup_id pointing to another product:

  • Recursively resolve missing IDs by following the chain.
  • Detect cycles and abort safely.
  • Filter out out-of-stock items from the final result.

Part 3 — Inventory Sync (Conflict Resolution)

Merge inventory updates from multiple warehouses with inconsistent field names; use the latest timestamp to win conflicts.


Reference Implementation

Retry With Exponential Backoff + Jitter

python
import random
import time
from typing import Callable, TypeVar

T = TypeVar("T")

class TransientError(Exception):
    pass

class PermanentError(Exception):
    pass

def retry(
    fn: Callable[[], T],
    *,
    max_attempts: int = 5,
    base_delay: float = 0.1,
    max_delay: float = 10.0,
    deadline_s: float = 30.0,
) -> T:
    start = time.monotonic()
    for attempt in range(max_attempts):
        try:
            return fn()
        except PermanentError:
            # 4xx — don't retry.
            raise
        except TransientError:
            # 5xx / timeout — back off and try again.
            if attempt == max_attempts - 1:
                raise
            elapsed = time.monotonic() - start
            if elapsed >= deadline_s:
                raise
            # Exponential backoff with full jitter.
            sleep_for = min(max_delay, base_delay * (2 ** attempt))
            sleep_for = random.uniform(0, sleep_for)
            # Respect deadline.
            sleep_for = min(sleep_for, deadline_s - elapsed)
            time.sleep(sleep_for)
    raise RuntimeError("unreachable")

Why full jitter (random.uniform(0, cap)) rather than fixed backoff? When many clients fail simultaneously (e.g. the server restarts), identical backoff schedules cause synchronized retry storms. Jitter spreads the retries out. AWS's own guidance recommends "full jitter" over "equal jitter" for this reason.

Cycle-Safe Product Resolution

python
def resolve_catalog(products_by_id, products):
    """
    Each product may be {"id": None, "backup_id": X, "in_stock": bool, ...}
    Resolve missing IDs by chasing backup_id; skip on cycles; filter out-of-stock.
    """
    resolved = []
    for product in products:
        resolved_id = _resolve_id(product, products_by_id)
        if resolved_id is None:
            continue  # Broken chain or cycle — drop.
        canonical = products_by_id[resolved_id]
        if not canonical.get("in_stock", True):
            continue  # Filter out-of-stock.
        resolved.append({**product, "id": resolved_id})
    return resolved

def _resolve_id(product, products_by_id):
    if product.get("id") is not None:
        return product["id"]
    # Floyd's / seen-set to detect cycles.
    seen = set()
    current = product
    while current.get("id") is None:
        backup = current.get("backup_id")
        if backup is None or backup in seen:
            return None  # Dead end or cycle.
        seen.add(backup)
        current = products_by_id.get(backup)
        if current is None:
            return None  # Dangling reference.
    return current["id"]

Inventory Sync (Schema Normalization + Timestamp Winner)

python
FIELD_ALIASES = {
    "sku_id": "sku", "product_id": "sku",
    "qty": "quantity", "count": "quantity",
    "updated_at": "timestamp", "ts": "timestamp",
}

def normalize(record):
    return {FIELD_ALIASES.get(k, k): v for k, v in record.items()}

def sync_inventory(*warehouse_feeds):
    winners = {}  # sku -> record
    for feed in warehouse_feeds:
        for raw in feed:
            rec = normalize(raw)
            sku = rec["sku"]
            if sku not in winners or rec["timestamp"] > winners[sku]["timestamp"]:
                winners[sku] = rec
    return list(winners.values())

Key Design Considerations

  • Classify errors before you retry. A blanket except Exception that retries on anything is the most common mistake — it retries on ValueError from bad input, which will never succeed. Classify into transient vs permanent first.
  • Respect deadlines, not just attempt counts. If each retry has a 30-second timeout and you allow 5 attempts, the worst case is 2.5 minutes. Caller deadlines matter.
  • Jitter > fixed delay. AWS architecture blog's "Exponential Backoff And Jitter" is the canonical reference. Full jitter (randomize across [0, cap]) is the default.
  • Cycle detection with a seen-set is O(n) per chain and simple. Floyd's tortoise-and-hare works too but is rarely worth the conceptual overhead for a bounded-depth chain.
  • Dangling references are distinct from cycles — the chain points at a product that doesn't exist. Both cases should fail gracefully, not raise.
  • Out-of-stock filtering happens AFTER resolution. A product may resolve to an ID whose canonical entry is out of stock — filter at that point, not on the raw record.

Variant (independent report — "Product Catalog Fallbacks")

A separate report describes the same three-part screen with different specifics for Parts 2 and 3.

Part 1 — API retries

A customer provides access to a mock Shopify-style API listing its product catalog. The API sometimes fails, and a first version of product fetching already exists. Implement retries so products are fetched more reliably. No other data processing is required in this part.

Part 2 — Fallback-chain resolution (relatedItems)

Each product may contain a fallbackSku pointing to a similar product. A fallback can point to another fallback, forming a chain.

For each product, add a relatedItems key containing every product reachable by following the fallbackSku trail. If the product has no fallback, relatedItems should be an empty array.

Part 3 — Product clustering (connected components)

Now treat each fallbackSku link as bidirectional. Products connected directly, indirectly, or in reverse belong to the same group.

For every product, update relatedItems to contain every other product ID in the same connected component, excluding the product itself.

Note from the report: the screenshots did not show the complete API implementation, product schema, retry policy, output-order requirement, or cycle-handling requirements for part two.

Environment (from CoderPad screenshots, March 2026 report)

Live CoderPad session (app.coderpad.io), Python, starter code split across src/api.py and src/main.py. The visible instruction headers were "Instructions Part 1" (retries against a mock Shopify-style API), "Part 2: Fallback Chain Resolution", and "Part 3: Product Clustering".

The candidate's in-progress solution visible in the screenshots (incomplete — shown for flavor of expected working style, not as a model answer):

python
def main():
    # in each index there is a product[x]["fallbackSku"] which has a id
    #   which is another product[x]["id"]
    # create a new key value pair in each product[x] called
    #   "related_items" : [every id which continues in fallbackSku]

    # goal #1: create a dictionary key id : index in product aray

    idmapping = dict()
    for i in range(len(products)):
        dictionary = products[i]
        idmapping[dictionary["id"]] = i

    for i in range(len(products)):
        products[i]["related_items"] = []
        temp = products[i]
        prev = set()
        while temp["fallbackSku"] and temp["id"] not in prev:  # while it has a next pointer
            prev.add(products[i]["id"])
            temp = products[idmapping[temp["fallbackSku"]]]
            products[i]["related_items"].append(temp["id"])

    for i in range(len(products)):
        print(products[i]["id"])
        print(products[i]["related_items"])

main()

Takeaway: Part 2 reduces to a guarded linked-list walk over fallbackSku (seen-set for cycles); Part 3 is connected components over the undirected link graph (union-find or BFS).

Source: community report (CoderPad screenshots), March 2026


Follow-up Questions

  1. How would you distinguish transient from permanent errors when the API just returns 500 for everything? (Inspect response body / error code for retryability hints; fall back to treating all 5xx as transient.)
  2. What if two warehouse feeds have the same timestamp for the same SKU? Pick a deterministic tiebreaker — warehouse ID, record hash, etc. Non-determinism is the bug.
  3. How do you make the retry function async-friendly? Replace time.sleep with asyncio.sleep; accept an async callable; wrap with asyncio.wait_for for per-attempt timeouts.
  4. Could the product resolution be cached? Yes — memoize _resolve_id on product dict identity or by backup_id chain head. Avoid caching when the catalog is mutable without invalidation hooks.