Technology Stack
Python, scikit-learn, VnCoreNLP, Streamlit, NLP, Feature Engineering
Overview
Duration: 10/2021
A machine learning system that classifies Vietnamese news articles as genuine or fabricated, taken from raw text through to a deployed, usable interface.
Vietnamese makes this harder than the equivalent English task in one specific and important way: it is not whitespace-tokenised. A Vietnamese "word" is frequently multiple syllables separated by spaces — hoc sinh is one word meaning "student", not two words. Splitting on whitespace, the default in almost every NLP tutorial, destroys the semantic units the model needs before training even begins. Handling that correctly is the difference between a model that works and one that appears to.
Preprocessing
The pipeline addresses Vietnamese-specific requirements rather than applying a generic recipe:
- Word segmentation through VnCoreNLP, so multi-syllable words are treated as single tokens — the single most consequential step in the pipeline
- Normalisation — case folding and Unicode normalisation, which matters in Vietnamese where the same diacritic can be encoded multiple ways and would otherwise produce distinct tokens for identical words
- Stopword removal using a Vietnamese-specific list, since English stopword lists are useless here
- Noise removal — HTML tags, special symbols, and punctuation stripped from scraped article text
- Feature extraction — the cleaned token stream converted to numeric vectors, since the models require numerical input
Exploratory Analysis
Before modelling: checking for missing and malformed records, measuring class balance — critical here, because an imbalanced dataset makes accuracy a misleading metric and a model that always predicts the majority class look successful — and profiling text statistics including article length distribution across both classes.
Modelling
Multiple classifiers were trained on the preprocessed corpus and compared rather than settling on the first one that produced a number. Both linear and non-linear approaches were evaluated, with the trained models serialised for use at inference time in the deployed application.
The dataset is the VFND Vietnamese fake news corpus — 223 labelled articles across two classes. That is a small corpus, and the honest conclusion from working with it is that the preprocessing pipeline and the evaluation methodology are where the real work is; absolute accuracy on 223 samples is not a claim worth making. The pipeline is what would scale to a larger corpus unchanged.
Data source: VFND Vietnamese Fake News Dataset
Deployment
The trained models are served behind a Streamlit interface: paste an article, select which model to run, and get a classification back. Deployed publicly rather than left as a notebook — because a model that only runs on the machine it was trained on has not really been delivered.
This project covers the full data science cycle on a language most tooling is not designed for: understanding why the standard preprocessing pipeline fails on Vietnamese, replacing the parts that need replacing, comparing models honestly, and shipping the result as something someone can actually use.