Session 13Session note

Cardputer and a Kurdish Ardalani dataset

Sirwan, Omid, Pooria, Saman, Arash, and Foad discuss recognising Kurdish numbers on Cardputer and building a text and audio dataset for Kurdish Ardalani.

Participants: Sirwan, Omid, Pooria, Saman, Arash, Foad

We started with a few general discussions related to Iran before moving to the main technical topics.

Cardputer Experiment

Sirwan shared a quick update on his Cardputer experiment, where he trained a small model to recognise the Kurdish numbers 1, 2, and 3.

A Kurdish Ardalani Dataset

We discussed the idea of fine-tuning a language model for Kurdish Ardalani. The main challenge is the lack of suitable training data, so we agreed that we should start building the dataset ourselves.

One approach is to record sample dialogues together with their written transcripts. Sirwan also suggested creating a smaller set of high-quality recordings first, then experimenting with audio models to generate variations of the same speech. This could help us expand the dataset without requiring many different people to record the same material.

Next Steps

  • Find a suitable written Kurdish Ardalani book.
  • Contact the author and ask whether we can get access to the book’s digital text.
  • Use the text as the basis for creating our initial text and audio dataset.
  • Start recording samples and experiment with generating additional audio variations.
All sessions