The Journey Behind My Handwriting Generation Pipeline
Warning: This post is long, messy, and built on curiosity. But it’s real this is how I built my first handwritten generation pipeline from scratch.
Hey guys, it's Manikanta. I'm genuinely proud to share my recent publication on Synthetic Handwritten Generation, a project I can truly call my own from start to finish.
This whole idea was born from a universal student dream: escaping the drudgery of writing hectic assignments. I was traveling home on the university bus one day when a friend complained, "I'm done with writing these assignments." We laughed, but it sparked a question in my mind: why hasn't anyone built an app that takes a sample of your handwriting and generates text in that same natural style?
Driven by curiosity, I dove into the research on ScienceDirect, looking at papers on OCR (Optical Character Recognition). Most focused on extracting and identifying text, but I needed to generate it. My past experience with GANs, which create similar-looking images, and Auto-Encoders, which extract outlines, felt like pieces of a puzzle. I immediately texted the idea to myself so I wouldn't forget.
My first attempt was straightforward: I wrote out A-Z, a-z, and 0-9 on a paper and tried to extract each character individually. It was a complete failure. Tools like Tesseract-OCR and Easy-OCR kept grabbing entire rows or nothing at all. Frustrated, I had to step away from the problem for a while.
The breakthrough came unexpectedly. My professor asked for some old code, and while reviewing my files, two concepts suddenly connected in my mind: Edge Detection and YOLO. Edge Detection highlights the outlines of an image, and YOLO places bounding boxes around objects to identify them. My interest took a complete U-turn.
I immediately applied edge detection to my character sheet. Boom the edges of every character were perfectly highlighted. The next logical step was to draw bounding boxes around those edges. Another boom! The technique successfully isolated all 62 characters. I had my foundation.
Now, the real architectural work began on the pipeline. My initial plan was: Edge Detection + Bounding Boxes -> CNN -> U-Net -> GANs -> Rendering. But I quickly realized that handling a multi-class detection problem was better suited for YOLO, so I changed it to: Edge Detection + Bounding Boxes -> YOLO -> U-Net -> GANs -> Rendering.
Then I overthought the U-Net's role. U-Net is for segmentation, focusing on the core object and removing the background. I worried that segmenting a character before the GAN would leave me with only an outer edge, making generation too hard. So, for the third time, I reshuffled the pipeline to its final form: Edge Detection + Bounding Boxes -> YOLO -> GANs -> U-NET -> Rendering.
With the blueprint set, I started collecting data. I gathered 93 handwritten samples from my friends—the best I could manage. Annotating all 93 images in Roboflow was the most grueling part; it took 3-4 days, and I owe a huge thanks to my sister for her help. I trained a YOLOv8s model, but the results were disastrous, with a Mean Average Precision (mAP) barely above zero. A newborn baby could have done better!
I used Roboflow's augmentation feature to expand my dataset to 243 images, but the mAP only climbed to 30-40. Staring at the confusing confusion matrix, I had a new idea: why not extract every single character using the labels I already had and create individual class folders? It worked. I resized the 27,000 resulting images to 64x64 and reformatted them for YOLO, effectively treating each full image as a bounding box. This time, training was a success, finishing with a solid mAP-95 of 0.75. The first major task was complete.
From here, the path felt smoother. The next step was the GAN, to generate new characters that looked natural. With nine types of GANs to choose from, I fell into a dilemma. I randomly started with StyleGAN, famous for generating realistic human faces. Unsurprisingly, it started generating faces instead of characters! My hope began to fade.
Further research led me to Conditional GANs, which use class labels to guide the generation. Perfect! When YOLO detected a character, it would be sorted into a class-labeled folder to train the GAN. But a new problem emerged: similar-looking characters like (0, o, O), (5, s, S), and (2, Z) were being grouped together, reducing our 62 classes to just 30-40 unique folders. Instead of fighting it, I decided to embrace these 19 clusters, allowing the model to substitute a detected 's' for a '5', for instance.
The final hurdle was achieving a natural, flowing handwriting style. Simply rotating characters created ugly black padding. The solution came from an Evolutionary Algorithm called CMA-ES (Covariance Matrix Adaptation-Evaluation Strategy), which intelligently shifts pixels within the matrix to create variation without adding noise. It worked beautifully.
The U-Net phase was surprisingly straightforward after all that, and soon I was rendering the final text. To be completely honest, I wasn't impressed with the final output. The metrics told the story: my FID score (which measures how real generated images look) was 240, indicating they were "super fake" compared to the ideal score of under 20. My segmentation scores like Inception Score and DiceID were around 50, far from the target of 80. The rendered characters just weren't as clear as I'd hoped.
Despite the imperfect results, I've documented the entire journey and submitted it to journals. This is the story of how I built my first complete pipeline idea, with all its twists, turns, and late-night breakthroughs.
I know research is all about inventing and trying something new, but here, everything I used has been applied in many other applications just not in the way I used it. This is only the beginning of my research journey. After this, the topic gained more add-ons and eventually transformed into a quality paper. Stay tuned for more project workings; I’ll update you on what happens next. Have a good day. ✌️