← Datasets
Image + TextText recognition (calligraphy)2024
ViCalligraphy
Vietnamese calligraphy (thư pháp) word recognition
15,541
Word crops
12,432 / 3,109
Train / Test
Word-level
Annotation
Vietnamese
Language
Samples
About
ViCalligraphy is a word-level recognition dataset of Vietnamese calligraphy (thư pháp) — brush-written, highly stylized glyphs with flowing strokes, ligatures and ornamental flourishes that depart sharply from printed or everyday handwritten fonts. It contains 15,541 cropped word images (12,432 train / 3,109 test), each annotated with its Vietnamese transcription including full diacritics. The dataset targets a domain underserved by existing scene-text and handwriting benchmarks, where artistic brush strokes and cursive connections make character boundaries ambiguous. The samples below are a small preview, and the full dataset is available on request.