The Telegraph Sorter
Reading the wires and sorting every message.
News Topic ClassificationNLP Architecture Comparison
91.9%accuracy, 0.919 macro F1
Reading the wires and sorting every message.
News Topic ClassificationNLP Architecture Comparison
91.9%accuracy, 0.919 macro F1
A controlled comparison, run by Araf as sole implementer, of how far preprocessing, word representation and architecture each move a news topic classifier.
| Axis | What was compared |
|---|---|
| Preprocessing | Three preprocessing strategies. |
| Representation | Two word representations: TF-IDF and a custom Skip-gram. |
| Architecture | Six architectures: DNN, SimpleRNN, GRU, LSTM, Bi-GRU and Bi-LSTM. |
Exploratory analysis drove every cleaning step: bigram frequency isolated HTML noise, POS tagging justified lemmatisation over stemming, and 22,067 duplicate records were removed.
Two recurrent baselines sat at chance accuracy. The diagnosis was exploding-gradient collapse; gradient clipping and targeted regularisation brought them back.
| Baseline | At chance | Recovered |
|---|---|---|
| LSTM | 25.0% | 91.4% |
| SimpleRNN | 26.6% | 87.3% |
91.9%
accuracy, best configuration
0.919
macro F1
22,067
duplicate records removed
There is no public repository or notebook for this project; this page is the account.
Python, TensorFlow/Keras, Gensim, scikit-learn