Fine-Tuned LLM: Messy Text to Clean JSON by Rishi BarapatreFine-Tuned LLM: Messy Text to Clean JSON by Rishi Barapatre

Fine-Tuned LLM: Messy Text to Clean JSON

Rishi Barapatre

Rishi Barapatre

Custom Fine-Tuned LLM: Messy Text to Clean, Structured Data

I fine-tuned a small open-source language model (Qwen3-4B) to read unstructured text and return clean, machine-readable JSON. It improved exact-match accuracy from 61.8% to 73.4% on held-out test data, trained on a free GPU.
Demo domain: medical case reports (pulling the drug name and the adverse effect out of a sentence). The same approach works for invoices, resumes, support tickets, legal clauses, product listings, or any text you need turned into structured fields.

The problem

Valuable information is buried in free-form text. Typing it into spreadsheets by hand is slow and error-prone, and general-purpose chatbots often give inconsistent or invented answers.

What it does

Input:
"Intravenous azithromycin-induced ototoxicity was reported in a patient receiving high-dose therapy."
Output:
{
"drug": "azithromycin",
"effect": "ototoxicity"
}

Results (500 unseen test sentences)

Metric Base model Fine-tuned Change Both fields exactly correct 61.8% 73.4% +11.6 points Field-level F1 85.9% 90.2% +4.3 points Drug name F1 90.5% 94.0% +3.5 points Valid JSON output 100% 100% no change
Test sentences were kept completely separate from training data. Results on your own data will differ and depend on data quality and volume.

Why it's practical

Low cost. Trained on just 3,000 examples using a free GPU (Google Colab T4).
Small and portable. The fine-tuned adapter is only about 66 MB on top of the base model.
Consistent format. Output is clean JSON your systems can use directly.
Reproducible. Full training and evaluation notebooks are included.

Where this fits

Extracting fields from documents, forms, or emails
Turning customer messages or tickets into structured records
Cleaning scanned or OCR text into spreadsheet-ready data
Domain-specific extraction where generic prompts are inconsistent

Built with

Python, PyTorch, Hugging Face Transformers, PEFT/LoRA, TRL, bitsandbytes (4-bit quantization), Google Colab

Want a custom model for your data?

I fine-tune and deploy extraction models and AI assistants for small businesses. Contact: rishibarapatre@gmail.com

Like this project

Posted Oct 9, 2026

Fine-tuned Qwen3-4B with QLoRA to turn messy clinical text into clean JSON. Exact match rose from 61.8% to 73.4% on a free GPU.