The Telegraph Sorter

Reading the wires and sorting every message.

News Topic ClassificationNLP Architecture Comparison

91.9%accuracy, 0.919 macro F1

Claimed

The Job

A controlled comparison, run by Araf as sole implementer, of how far preprocessing, word representation and architecture each move a news topic classifier.

How it was done

The grid

AxisWhat was compared
PreprocessingThree preprocessing strategies.
RepresentationTwo word representations: TF-IDF and a custom Skip-gram.
ArchitectureSix architectures: DNN, SimpleRNN, GRU, LSTM, Bi-GRU and Bi-LSTM.

Preprocessing from the data, not from habit

Exploratory analysis drove every cleaning step: bigram frequency isolated HTML noise, POS tagging justified lemmatisation over stemming, and 22,067 duplicate records were removed.

Two dead baselines, brought back

Two recurrent baselines sat at chance accuracy. The diagnosis was exploding-gradient collapse; gradient clipping and targeted regularisation brought them back.

BaselineAt chanceRecovered
LSTM25.0%91.4%
SimpleRNN26.6%87.3%

The Take

There is no public repository or notebook for this project; this page is the account.

The stack

Python, TensorFlow/Keras, Gensim, scikit-learn