We believe in trusting relationships with our stakeholders. That’s why we’re sharing our recipe for success (without giving away the secret sauce).
What we predict
For every borrower in a partner NBFC’s book, our model estimates the probability that they fall 30, 60 or 90+ days behind on repayment. Our model is capable of predicting at multiple time horizons, so we can price an entire loan pool at once.
Our Model is our Moat
A standard credit model learns from finished stories: loans that either repaid or defaulted. But at any moment, a large share of a lender’s portfolio is still outstanding. A two-year loan disbursed eight months ago hasn’t succeeded or failed yet. Statisticians call this censored data.
The simple fix is to train only on loans which are old enough to have a final outcome. But that discards a large share of the most recent data. We found that this is the share that matters most. Repayment capacity moves with current economic conditions: crop prices, climate, local labour demand and funding conditions across the sector. A model that doesn’t learn from such recent data can’t successfully predict current borrower’s repayment capabilities.
How the model works
Picture checking in on every loan once a month and asking one simple question: did this borrower fall behind this month, or are they still on track? Our model learns from all of those monthly check-ins.
Take a loan that has been paid on time for eight months and isn’t finished yet. Other models would ignore it until it ends. Ours learns from what we already know: eight good months. Nothing gets thrown away and nothing gets made up.
Put the monthly chances together and you get the chance that a borrower falls behind within three months, six months or a year.
To learn these chances, the model builds lots of small decision trees. Each tree is a chain of yes-or-no questions, and each new tree fixes the mistakes of the ones before it. That lets the model spot patterns that only show up in combination: a missed payment can mean something very different in the rainy season than at harvest time, or in one village compared with the next.
What goes in
Every prediction draws on five layers of information, all measured as of the moment the prediction is made, so the model never sees the future it’s predicting.
Repayment history
How the borrower has paid so far: timing, gaps, partial payments, previous loans.
Credit bureau records
Score, number of active loans and delinquencies with other lenders. Whether a borrower has a bureau record at all is itself informative, and the model treats it that way.
Borrower and loan details
Demographics, loan size, product, term and repayment frequency.
Geography and season
Dynamic environmental data: where the borrower lives and what is happening there, month by month, from agricultural seasons to regional risk.
Social environment
Microfinance runs on relationships: loan officers visit borrowers at home, and borrowers repay in groups that meet weekly. We measure the track record of each borrower’s officer and group, with statistical shrinkage so small groups aren’t judged on a handful of loans.
How we check ourselves
We validated our model on nearly half a million historical microloans from established Indian microfinance lenders. The model's performance is measured by AUC: the probability that, given a borrower who is delinquent and one who isn't, the model successfully ranks the delinquent borrower as riskier. 0.5 is a coin flip; 1.0 is perfect.
High scores in credit modelling are easy to fake by accident. Information about the outcome can slip into the inputs, or the test set can quietly resemble the training set. We run three checks on every version of the model:
Borrower-level splits. Many borrowers take several loans. We test with all of a borrower’s loans held out together, so the model can’t recognise a borrower it has already seen. The score drops by just 0.015.
Time-based splits. We train on older loans and test only on newer ones, the way the model is actually used. This is where the 0.83 comes from. The drop is concentrated in the shortest horizons, where new loans simply haven’t had time to show any history. That is the pattern of immature data, not of a model that has been cheating.
Label shuffling. We randomly reassign every loan’s outcome to a different loan and retrain. If the inputs secretly encoded the answer, the model would still score well. It doesn’t: what little it retains is fully explained by the age of the loan, a known and intended input.
One problem is smaller in microfinance than in other lending. Credit models only ever see the borrowers who were approved. In credit cards, that can be a small fraction of applicants; in microfinance, most applicants are approved, because groups and officers screen them before a formal application exists. The borrowers we learn from look much like the borrowers we score.
What it means for our partners
When a partner NBFC opens its book to us, we score every account, month by month. We select the loans we are most confident in, build them into a securitised pool, and price that pool on our own expected losses. The partner gets cash for the loans they’ve already made. We get a portfolio we understand borrower by borrower.
Figures are from our current development runs and will be updated as the model is retrained on new data.