Validating and correcting PII annotations in synthetically generated Dutch-language conversations, covering both nl-NL and nl-BE (Flemish) locales.
The work goes beyond simple labelling. Each conversation is reviewed for missing or incorrect entity spans, offset accuracy, and whether every instance of a PII entity is consistently tagged across turns. On top of that, the Dutch language quality is assessed — grammar, naturalness, locale-appropriateness, and conversation coherence — with corrections and explanations provided where needed.
PII entity types covered include names, addresses, phone numbers, email addresses, IBANs, ID numbers, dates, usernames, and passwords, each validated against the correct Dutch or Belgian format.
Quality targets: inter-annotator agreement above 95%, full locale validity, and self-QC accuracy of 98% or higher.
Native Dutch speaker covering both standard Dutch and Flemish — not two separate skill sets, but one annotator who knows the difference.
Like this project
Posted Sep 4, 2026
Validating and correcting PII annotations in synthetically generated Dutch-language conversations, covering both nl-NL and nl-BE (Flemish) locales.
The work ...