Ten thousand registrations and the number I was avoiding
Ten thousand registrations is a good number to say out loud. I said it a lot.
The number I wasn't saying was how many of those people opened the app in their second week. I knew roughly what it was, in the way you know a thing without looking at it directly, and I put off building the query for longer than I should have.
When I finally built it, it was bad. Not catastrophically bad for an education product, but bad enough that every plan I had for the next quarter was answering the wrong question. We were spending on acquisition to fill a bucket with a hole in it.
Retention in exam prep isn't retention
The first thing that went wrong was borrowing metrics from products that aren't like ours.
D7 and D30 are the standard, and for an exam-prep app they're close to meaningless. A student preparing for CSIR NET has a date. Everything about their behaviour bends around it. Usage climbs for four months and falls off a cliff the day after the exam, and that cliff is not churn, it's success. Meanwhile a student who registers in March for a December exam and doesn't come back until August hasn't lapsed. They were doing coursework.
So a flat retention curve was the wrong target. What we ended up watching instead was whether a student's activity was tracking their own exam date. Two students with identical D30 could be in completely different situations, and only one of them needed us to do something.
That reframing took an embarrassingly long time, and it came from talking to students rather than from any dashboard. The dashboards had been quietly telling me a story about a leaky consumer app, and we weren't one.
The funnel that mattered
Once we stopped looking at sessions and started looking at attempts, things got clearer.
The single strongest predictor of whether somebody was still with us three months later was whether they had completed a full-length mock test in their first fortnight. Not watched lectures. Not downloaded notes. Completed a test, seen a score, and looked at the solutions.
I want to be careful here, because this is the point where people usually announce they've found the growth lever and start forcing everyone through it. Correlation was doing a lot of work in that finding. Students who sit a three-hour mock in week two are, on average, more serious to begin with. Making a casual student take a test earlier doesn't turn them into a serious one.
What we could do was remove the friction between a serious student and their first test. Which turned out to be mostly boring: the test list defaulted to full syllabus tests they weren't ready for, so the sensible ones bounced off. Adding shorter chapter tests at the top of that list moved the number more than any feature we shipped that year.
What actually moved the needle
The best retention work we did wasn't a feature at all. It was a schedule.
We started running a weekly mock at a fixed time on Sunday morning, with results and an all-India rank published on Sunday evening. Same time every week. Nothing about it was technically interesting. Attendance built over about two months and then stayed, and the students who joined that rhythm behaved completely differently from the ones using the app whenever they felt like it.
A deadline you share with other people is a stronger product than most of the things we'd been building. That has been the most transferable lesson from Egxam into everything else I've worked on, and it isn't a software lesson.
The other lever was the teachers. Ten-odd people writing content, and the difference between an engaged one and a coasting one showed up in the numbers within a month. Not through anything sophisticated, just completion rates on their material and what students wrote in feedback. That was uncomfortable to look at as a manager, because it makes a conversation unavoidable that you'd otherwise let slide for a year.
The finding I didn't like
Our most carefully made content was not our most used content.
We had a set of concept videos, properly scripted, genuinely good, that took weeks to produce. They sat there. Meanwhile a set of quickly-made PDFs of previous years' questions with worked solutions got opened constantly, including by students who never touched anything else.
My first reaction was that students were optimising badly and needed guidance. My second, better, reaction was that they knew exactly what they were doing. Three months before an exam, a worked solution to a question that has actually been asked is more valuable than a beautiful explanation of a concept. We were producing what we were proud of rather than what was needed at that point in the cycle.
We didn't stop making the videos. We stopped making them in October.
The reviews said something different from the ratings
We sit at 4.4, which is fine and which tells you nothing. The written reviews were far more useful, and about a third of them were about a single thing: the app's behaviour on poor connections. Not content, not pricing, not features. Whether a downloaded test would survive a train journey.
None of that was visible in any product metric we had, because the students it affected most were the ones least likely to complete the flows we were measuring. Failure is silent in your analytics almost by definition.
I now read the one-star reviews first, and I've kept that habit at work. Support tickets and reviews are the only channel where people tell you about the thing your instrumentation couldn't see.
The second-week number, for what it's worth, roughly doubled over the following year. Most of that came from the Sunday mocks and the chapter tests, which between them cost less engineering time than one of the concept video series did.