Paper: Lightweight baselines for medical abstract classification: DistilBERT with cross-entropy as a strong default. J. Liu, L. Wang, S. Liu, X. Hu. 8th International Conference on Machine Learning and Natural Language Processing (MLNLP), 2025.

Code: github.com/JackyLiu47/medical_abstract

This post covers both the paper and the repository behind it, since they are the same project.

The question

When you need to classify medical text, it is tempting to reach for the biggest domain-specific model and a clever loss function. We asked the boring question first: how far does a small, general-purpose model with a standard loss get you?

The task is five-way classification on the TimSchopf/medical_abstracts dataset from Hugging Face.

The experiment grid

One training script, train_medabs.py, drives everything from the command line:

  • Models: BERT-base, DistilBERT, plus domain models BioBERT, SciBERT and PubMedBERT
  • Losses: cross-entropy, class-weighted, and focal loss
  • Extras: bf16/fp16 mixed precision, per-class threshold tuning, temperature-scaling calibration, label smoothing
  • Reproducibility: fixed seeds and stratified splits
  • Outputs: metrics JSON, training logs, confusion matrices, per-class charts and loss curves

What we found

We report two regimes. Raw results use plain argmax decoding, which isolates the effect of the loss function. Tuned results add post-hoc operating-point selection: temperature scaling on the validation set, then class-wise thresholds chosen on validation and frozen on test.

Dot plot of accuracy and macro-F1 for six raw configurations and three threshold-tuned DistilBERT configurations; tuned focal loss reaches 77.6% macro-F1.

Raw (argmax). DistilBERT with plain cross-entropy is the strongest and most stable setting (64.6% accuracy, 64.4% macro-F1). It matches BERT-base while being about 40% smaller and roughly twice as fast, and class-weighted and focal losses gave no gain here.

Tuned. Calibration and thresholds change the picture a lot. Macro-F1 rises to 70.7% for cross-entropy and to 77.6% for focal loss (82.2% accuracy), so under deployment-style tuning focal loss benefits most.

The practical takeaway of the paper: start with a compact encoder and cross-entropy, then add lightweight calibration and thresholding when deployment needs a better macro balance. Judge the training objective and the deployment operating point separately.

Figure from the paper: macro-F1 for six raw and three tuned configurations.

Left, grey: raw configurations. Right, blue: tuned configurations. Figure 1 of the paper.

Try it

pip install -r requirements.txt
python train_medabs.py --model_name distilbert-base-uncased --loss_type cross_entropy

Citation

@inproceedings{liu2025lightweight,
  title={Lightweight baselines for medical abstract classification: DistilBERT with cross-entropy as a strong default},
  author={Liu, Jiaqi and Wang, Lanruo and Liu, Su and Hu, Xin},
  booktitle={2025 8th International Conference on Machine Learning and Natural Language Processing (MLNLP 2025)},
  year={2025}
}