Back to Bloomberg questions
CodingSoftware Engineer

Bigram Autocomplete

Frequency: Reported


Create a program with two parts:

  1. A training section that accepts sequences of words and counts bigram frequency: how often each word is followed by every other word.
  2. An API that accepts a known word and returns the word most frequently observed immediately after it.

Sample data

python
trainingData = [
    ["I", "am", "Sam"],
    ["Sam", "I", "am"],
    ["Green", "Eggs", "I", "like"],
    ["Green", "Eggs", "and", "ham"],
]

After training on this data, predict("I") returns "am".

Use a class or function named AutoComplete to train the model, then implement predict(word).

Partial implementation visible in the source

python
from collections import defaultdict, Counter

class AutoComplete:
    def __init__(self):
        self.model = defaultdict(Counter)

    def train(self, trainingData):
        for sentence in trainingData:
            for i in range(len(sentence) - 1):
                w1, w2 = sentence[i], sentence[i + 1]

The screenshot ends at that line. It does not preserve the rest of train, the implementation of predict, behavior for an unknown word, or frequency tie-breaking.