Kurdish speech-to-text, AI code review, and cybersecurity
Saman, Omid, Foad, and Sirwan discuss Kurdish speech-to-text data, AI models and code review, frontier-model cybersecurity, and European jobs and living costs.
Participants: Saman, Omid, Foad, Sirwan
Kurdish Speech-to-Text
We discussed recording KurdSoftware sessions and using them to improve Kurdish speech-to-text.
Arash and I have already explored fine-tuning an NVIDIA speech-to-text model. The main limitation is the lack of high-quality Kurdish data, especially for the Ardalani dialect.
Recording sessions could help, but we still need accurate transcripts. One approach is to generate initial transcripts with an existing model and manually correct them.
Foad and Saman also suggested some Kurdish data resources and people to follow.
AI Models
Foad shared his positive experience with Qwen and DeepSeek for daily use.
We also discussed model distillation and how smaller models are becoming increasingly capable.
AI Code Review
We discussed human vs AI pull request reviews.
The main conclusion was that as AI takes over more implementation and review work, engineering responsibility and ownership become even more important.
Cybersecurity
We discussed the growing cybersecurity capabilities of frontier models.
I mentioned OpenAI’s Astra model, including its strong coding, maths, and cybersecurity capabilities and its demonstrations with Unreal Engine workflows.
We also discussed prompt injection, including hidden instructions in CVs designed to influence AI recruitment systems, as well as the Hugging Face incident involving GPT-5.6 Sol and another internal OpenAI research model.
Job Market and Cost of Living
We discussed recent European tech lay-offs, including the closure of a Stockholm office with around 450 employees.
We also compared the UK and Sweden, especially the high cost of commuting and train travel in the UK.