uLearn
An AI language-learning platform with real-time voice tutoring, adaptive flashcards, and grammar support across four languages. Designed, built, and operated solo.
Timeline
Live
Team
Solo developer
Role
Founder & Full-Stack Developer
Overview
uLearn is an AI language-learning platform for German, English, Russian, and French. A real-time voice tutor holds spoken conversations with contextual corrections, and whatever you read, including YouTube transcripts, turns straight into adaptive flashcards.
The Challenge
The challenge isn't any single provider, it's running five of them solo: AWS for hosting and the database, Groq for speech-to-text, OpenAI for text-to-speech, Gemini for chat and text generation, and Webshare as a proxy for the YouTube videos learners study from. One person keeps all of that integrated and running, with no team to split it across.
The Solution
The technique that resolves this challenge is a microservices architecture: each provider lives inside its own independent Go service instead of a shared codebase. That does two things: touching Gemini's chat generation means opening one specific repository, not a file buried inside a bigger project that touches a dozen other things, and an outage on one provider, say Groq's STT API, stays contained to that service instead of taking down chat or flashcards too.
Nine small services, each owning one job, deployed on AWS ECS and provisioned with Pulumi, keep both of those manageable for one person instead of a team.
A conversation, not a script
The AI tutor holds an actual back-and-forth conversation out loud. It reacts to what you say, corrects you in context, and builds a practice scenario on request instead of running through a fixed script.
That's a big enough engineering problem on its own that it got its own write-up: keeping a WebSocket open instead of round-tripping HTTP, gating the mic with real voice-activity detection, and making sure the assistant never talks over itself.
Read anything, turn it into review
Highlight a word or a phrase anywhere you're reading, including inside a YouTube transcript, and it becomes a flashcard on the spot: translation, grammar, and context already filled in. No manual data entry, no separate app for vocabulary. See it yourself in a guided tour of reading and looking up a word.
Review isn't a static deck either. Cards come back spaced by how well you already know each one, not the same order every time. Here's a guided tour of the review page too.
Four languages, on their own terms
Flashcards are tailored to each language's own grammar, not one template stretched across all four. Russian nouns get full declension tables, since Russian's case system is complex enough to warrant one. German nouns decline too, but simply enough that a dedicated table would be overkill, so its flashcards skip it. French leans on verb conjugation instead. See what that actually looks like a bit further down.
That's possible because each language has its own DynamoDB table, and DynamoDB is schemaless: a new field, like a grammar table a language didn't need before, can be added straight into production, no migration required.
What it actually looks like
A flashcard, filled in for you
Translation, meaning, real example sentences, and word frequency, generated from a single tap, not typed in by hand.
The conjugation table it builds
French leans on verb conjugation, so that's what its flashcards get: every tense, generated straight from the verb.
Grouped by grammar, not by date
Collections organize flashcards by what they teach, like every French verb that takes être as its auxiliary, not by when they were added.
Built and run by one person
None of that runs on a single server. AWS for hosting and the database, Groq for speech-to-text, OpenAI for text-to-speech, Gemini for chat and text generation, Webshare as a proxy for the YouTube videos learners study from: five outside providers, and one person keeping all of them integrated and running.
The pattern that makes that manageable is a microservices architecture. Nine independent services, each owning one job, deployed on AWS ECS. Touching Gemini's chat generation means opening one specific repository, not a file buried inside a bigger project that touches a dozen other things, and a hiccup in one provider, say Groq's STT API going down, stays contained to that one service instead of taking the whole app with it.
Where it runs
uLearn is live at ultimatelearning.app, supporting German, English, Russian, and French, on Free, Pro, and Premium tiers. The web frontend is served through HTML6, the same template engine built at Eldøy Projects, now running a product of its own in production.
Interested in this project?
I'd love to discuss the technical details, challenges faced, and lessons learned from building this project.
Get in Touch