6

Vietnamese Fake News Detection — NLP Classifier

A text classification system for Vietnamese news built end to end — language-specific preprocessing with proper word segmentation, feature extraction, multiple compared models, and a deployed web interface where an article can be pasted in and classified live.

Technology Stack

Python, scikit-learn, VnCoreNLP, Streamlit, NLP, Feature Engineering

Overview

Duration: 10/2021

A machine learning system that classifies Vietnamese news articles as genuine or fabricated, taken from raw text through to a deployed, usable interface.

Vietnamese makes this harder than the equivalent English task in one specific and important way: it is not whitespace-tokenised. A Vietnamese "word" is frequently multiple syllables separated by spaces — hoc sinh is one word meaning "student", not two words. Splitting on whitespace, the default in almost every NLP tutorial, destroys the semantic units the model needs before training even begins. Handling that correctly is the difference between a model that works and one that appears to.

Preprocessing

The pipeline addresses Vietnamese-specific requirements rather than applying a generic recipe:

Exploratory Analysis

Before modelling: checking for missing and malformed records, measuring class balance — critical here, because an imbalanced dataset makes accuracy a misleading metric and a model that always predicts the majority class look successful — and profiling text statistics including article length distribution across both classes.

Modelling

Multiple classifiers were trained on the preprocessed corpus and compared rather than settling on the first one that produced a number. Both linear and non-linear approaches were evaluated, with the trained models serialised for use at inference time in the deployed application.

The dataset is the VFND Vietnamese fake news corpus — 223 labelled articles across two classes. That is a small corpus, and the honest conclusion from working with it is that the preprocessing pipeline and the evaluation methodology are where the real work is; absolute accuracy on 223 samples is not a claim worth making. The pipeline is what would scale to a larger corpus unchanged.

Data source: VFND Vietnamese Fake News Dataset

Deployment

The trained models are served behind a Streamlit interface: paste an article, select which model to run, and get a classification back. Deployed publicly rather than left as a notebook — because a model that only runs on the machine it was trained on has not really been delivered.

This project covers the full data science cycle on a language most tooling is not designed for: understanding why the standard preprocessing pipeline fails on Vietnamese, replacing the parts that need replacing, comparing models honestly, and shipping the result as something someone can actually use.