Word2Vec & GloVe from Scratch
Built an end-to-end word-embedding pipeline from first principles: regex tokenization, frequency-cut vocabulary, Mikolov subsampling, an O(1) negative sampler from the unigram0.75 distribution, Skip-Gram, CBOW and GloVe with weighted least squares over sparse co-occurrence data.