Paper: Lightweight baselines for medical abstract classification: DistilBERT with cross-entropy as a strong default. J. Liu, L. Wang, S. Liu, X. Hu. 8th International Conference on Machine Learning and Natural Language Processing (MLNLP), 2025.
Code: github.com/JackyLiu47/medical_abstract
This post covers both the paper and the repository behind it, since they are the same project.
The question
When you need to classify medical text, it is tempting to reach for the biggest domain-specific model and a clever loss function. We asked the boring question first: how far does a small, general-purpose model with a standard loss get you?
The task is five-way classification on the TimSchopf/medical_abstracts dataset from Hugging Face.
The experiment grid
One training script, train_medabs.py, drives everything from the command line:
- Models: BERT-base, DistilBERT, plus domain models BioBERT, SciBERT and PubMedBERT
- Losses: cross-entropy, class-weighted, and focal loss
- Extras: bf16/fp16 mixed precision, per-class threshold tuning, temperature-scaling calibration, label smoothing
- Reproducibility: fixed seeds and stratified splits
- Outputs: metrics JSON, training logs, confusion matrices, per-class charts and loss curves
What we found
We report two regimes. Raw results use plain argmax decoding, which isolates the effect of the loss function. Tuned results add post-hoc operating-point selection: temperature scaling on the validation set, then class-wise thresholds chosen on validation and frozen on test.
Raw (argmax). DistilBERT with plain cross-entropy is the strongest and most stable setting (64.6% accuracy, 64.4% macro-F1). It matches BERT-base while being about 40% smaller and roughly twice as fast, and class-weighted and focal losses gave no gain here.
Tuned. Calibration and thresholds change the picture a lot. Macro-F1 rises to 70.7% for cross-entropy and to 77.6% for focal loss (82.2% accuracy), so under deployment-style tuning focal loss benefits most.
The practical takeaway of the paper: start with a compact encoder and cross-entropy, then add lightweight calibration and thresholding when deployment needs a better macro balance. Judge the training objective and the deployment operating point separately.

Left, grey: raw configurations. Right, blue: tuned configurations. Figure 1 of the paper.
Try it
pip install -r requirements.txt
python train_medabs.py --model_name distilbert-base-uncased --loss_type cross_entropy
Citation
@inproceedings{liu2025lightweight,
title={Lightweight baselines for medical abstract classification: DistilBERT with cross-entropy as a strong default},
author={Liu, Jiaqi and Wang, Lanruo and Liu, Su and Hu, Xin},
booktitle={2025 8th International Conference on Machine Learning and Natural Language Processing (MLNLP 2025)},
year={2025}
}
