Kurdish OCR, authentic data, and a first training experiment
Sirwan, Saman, and Arash discuss Asosoft OCR, AI-assisted cybersecurity, and a collection pipeline for authentic Ardalani data before a first one-epoch training experiment.
Participants: Sirwan, Saman, Arash
The meeting focused on Kurdish OCR, dataset collection, AI-assisted cybersecurity, and the first model-training experiment. The main decision was to prioritise authentic Ardalani Kurdish data over AI-generated expressions and build a dataset collection pipeline that can later support community contributions.
Asosoft — Kurdish OCR
Saman tested the Asosoft OCR tool using the Kurdish book he had previously shared in the Telegram group. The book contains many useful, mostly uncensored expressions in the Ardalani dialect, and the initial OCR results were promising.
Saman contacted the Asosoft team, who agreed to follow up about possible API access. We also discussed acknowledging Asosoft when publishing the model in the future.
Frontier Models and Cybersecurity
We discussed how frontier AI models can support penetration testing and vulnerability discovery. We shared a recent experience demonstrating how these models can identify potential security weaknesses in websites. The discussion highlighted their usefulness for security research and authorised testing.
Dataset Collection Pipeline
After Arash joined, we reviewed the Asosoft update and discussed building our own dataset collection process. The proposed approach was to gather authentic Kurdish text ourselves and build a pipeline to collect, clean, and prepare training data.
We also discussed exploring Telegram bots that let community members search the dataset and contribute expressions, so the dataset can expand incrementally through community contributions.
Authentic Data vs AI-Generated Expressions
Sirwan suggested generating 100 Kurdish expressions and using static mappings to convert them into Ardalani. Arash raised concerns about their quality, noting that many AI-generated expressions were inaccurate or unnatural.
Decision: Use authentic Kurdish expressions collected from real sources rather than synthetic expressions of uncertain quality.
First Training Experiment
We agreed to proceed with an initial experiment using the existing dataset. The goal is to test the training process, evaluate the results, and identify improvements before investing in a larger dataset. The first run will use one epoch.