DataData in FintechBankingFintech

How Tala built a credit engine for the world's most invisible borrowers

Tala lends to borrowers who don't exist in any credit bureau, using smartphone data as a substitute for a credit file. Here is what their model actually does, what it has produced, and what fintech data leaders can reasonably take from it.

Listen to the podcast

5 min

Chapters

Key takeaways

  • Treat "no credit history" as a format problem, not an absence of data, and go looking for behavioral signals elsewhere.
  • Build the advantage in the data pipeline, not the algorithm, because proprietary data beats a marginally smarter model.
  • Retrain constantly and assume a signal that worked in Nairobi in 2022 may be noise by 2026.
  • Run automated freshness and quality tests on input data, or accept that your scoring is guesswork with a green dashboard.
  • Pick your three most important data sources and answer how you would know within a week if they went stale.
Read the full transcript

Host:Welcome to the Leaders Insights podcast. Today's episode, how Tala built a credit engine for the world's most invisible borrowers.

Expert:There are roughly 2 billion adults on this planet with no credit history at all. Tala looked at that and said, we'll lend to them anyway.

Host:How is that not just charity that goes bankrupt?

Expert:Because charity doesn't return capital and Tala does. The trick everyone misses is that no credit history doesn't mean no data. It means no data in the format lenders got lazy about relying on. The bureau file. The tidy score. Tala went looking for signals somewhere else. The phone.

Host:The phone. So they're snooping through people's text messages to decide if they're creditworthy?

Expert:With permission, yes. And less creepily than it sounds. When you install the Tala app, you consent to share device data. How many contacts you keep, whether you top up your airtime before it runs out. How regularly you move money around. None of that is a moral judgment. It's a behavioral proxy.

Host:Yeti?

Expert:A stand-in signal for the stable financial habits a bureau score would normally capture.

Host:Come on, my contact list predicts whether I'll repay a loan? That sounds like the kind of thing that gets you sued.

Expert:It's not one variable. It's the pattern across hundreds. And here's the contrarian bit. The consensus in fintech is that you need better models. Wrong. Tala's early advantage wasn't a fancier algorithm. It was owning a data source nobody else bothered to collect. The model is almost boring. The data pipeline is the moat.

Host:That's a convenient thing for a data person to say.

Expert:It is, and I'll back it. Tala has dispersed over $6 billion across markets like Kenya, the Philippines, Mexico, India, to people who'd be an automatic decline anywhere else. You don't get there on model cleverness. You get there because you're the only one holding the raw material. MIT Sloan Management Review has made this point for years. In machine learning, proprietary data beats a marginally smarter algorithm almost every time.

Host:Fine, but smartphone behavior changes constantly. People switch phones, apps get updated, habits shift. Get your entire signal rod underneath you?

Expert:That's the real risk, and it's where most people underestimate the work. This is model drift. When the world moves and your predictions quietly get worse, while the dashboard still looks green. Tala has to retrain constantly because a signal that worked in Nairobi in 2022 might be noise by 2026. The lending decision is the easy part. Keeping the pipeline honest as behavior mutates. It's the actual job.

Host:How do you even measure whether the data's still good? Everyone claims their pipeline is clean.

Expert:You test it, ruthlessly. DBT Labs, and I'll flag they sell the data transformation tooling, so crosscheck against O'Reilly's independent surveys, publishes numbers on how many teams still don't run automated tests on their data. It's uncomfortably high. O'Reilly's data engineering work backs the direction. Research failures aren't the model. They're broken inputs nobody checked. A lender running on phone data with no freshness tests isn't doing analytics. It's doing astrology.

Host:So what stops a copycat from just cloning this? If the mode is data, someone with a bigger balance sheet buys their way in.

Expert:They try, but the data compounds. Every lone Tala issues and collects teaches the model something a newcomer doesn't know yet. That's the part balance sheets can't shortcut. You can buy servers. You can't buy repayment history you don't have. The invisible borrower stopped being invisible the moment Tala started keeping receipts.

Host:Give me the uncomfortable limit. Where does this model genuinely fail?

Expert:When the proxy stops meaning what you think. If everyone games airtime top-ups because they've learned it boosts their score, your signal is now theater. Behavioral data is only predictive while people aren't performing for the algorithm. The day they are, you're back to lending blind. Wando, Asa Kasti, and possibly to the most sophisticated fraudsters, not the honest poor.

Host:So what's the one thing a FinTech data leader should actually do with all this?

Expert:Stop obsessing over the model and audit your inputs first. Pick your three most important data sources and ask one question. How would I know this week if they went stale? If you can't answer, that's your project, not a bigger algorithm.

Host:Use for today's episode. O'Reilly Radar, Katie Nuggets, The New Stack, Towards Data Science, MIT Sloan Management Review, DPT Labs, Vendor, Data Tooling. Done. Want to know where you actually stand? Take the CDO self-assessment at MBA-training.com.

When Shivani Siroya founded Tala in 2011, she was trying to solve a problem that conventional credit scoring had simply declared unsolvable: how do you underwrite someone who has never had a bank account, a utility contract in their name, or a loan of any kind? The thin-file customer, in a mature market like the US or UK, is an edge case. In Kenya, Tanzania, the Philippines, and India, that description fits the majority of the adult population. Traditional FICO-style models need a minimum history to generate any score at all. Tala's first market, Kenya, had mobile-money penetration above 70 percent but formal credit bureau coverage well below that. The addressable population was enormous; the underwriting data, by conventional definitions, was absent.

How does Tala score borrowers with no credit history?

Tala's approach starts from a simple technical premise: a smartphone is a transaction ledger, a behavioral diary, and a social-network map, all at once. With explicit user consent obtained at onboarding, the Tala app requests access to a specific set of on-device signals: SMS transaction records from M-Pesa, call logs, app usage patterns, and metadata about contacts (not content, but frequency and reciprocity of communication). The company then runs these raw signals through a proprietary model to generate a credit score where none previously existed.

The feature engineering is where the real work sits. Tala's data scientists found that the ratio of incoming to outgoing M-Pesa transactions was predictive of repayment behavior, as was the diversity of counterparties a borrower transacted with. Someone who receives money from many different people, rather than a single employer, often has a more resilient informal income stream than a narrow transaction history would suggest. The stability of a contact network over time, meaning low churn in who a borrower calls, also correlated with reliability. None of these variables appear in any credit bureau.

This is not just feature selection; it is an entirely different theory of creditworthiness. Where bureau-based models measure a borrower's past behavior with formal financial institutions, Tala's model measures economic participation in the informal economy.Understanding what payment flows and contact patterns actually signal about a borrower's financial life requires building that analytical capability from scratch, because no vendor can sell you a pre-packaged proxy for M-Pesa reciprocity ratios.

Tala also invested heavily in model governance from an early stage, and for good reason. Lending to underserved populations in emerging markets triggers scrutiny from multiple directions simultaneously: consumer protection regulators in each operating country, global investors wanting evidence that loan pricing is defensible, and civil-society groups alert to any pattern that looks like it targets the vulnerable. The company publishes what its models consider and what they do not, and it explicitly excludes variables that could function as proxies for race, religion, or ethnicity under local fair-lending frameworks.Getting this wrong, even unintentionally, creates the kind of adverse action liability that can stop a lending operation entirely.

Tala's loan volume, borrower count and default rates

By 2026, Tala has disbursed more than 9 billion dollars in loans across its markets, according to company-published figures. It claims a customer base of over 8 million borrowers, the majority of whom had no prior formal credit history. Those numbers come from Tala itself, so treat them as indicative rather than independently verified.

What is verifiable from external reporting is the default rate benchmark. Tala has consistently cited default rates in the 7 to 10 percent range for its emerging-market portfolios, which compares favorably with informal moneylender losses in the same markets, though it is higher than the sub-3 percent defaults a well-collateralized prime lender would accept. The comparison that matters is not against a prime portfolio; it is against the previous option available to these borrowers, which was no formal credit at all, or a loan shark charging 200 percent annualized.

The model has also improved over time as the dataset has grown. Repayment behavior from earlier cohorts feeds back into the scoring model, which means Tala's competitive position compounds in a way that a new entrant without that behavioral history cannot easily replicate. This is the genuine moat in alternative-data underwriting: the data flywheel, not the algorithm.

What transfers from Tala to regulated Western lenders

For a CDO at a fintech operating in more regulated Western markets, the Tala case teaches several things that do not require a Nairobi office.

First, the principle that transactional behavior reveals creditworthiness generalizes, but the specific signals do not. In a UK or EU context, open banking access under PSD2 gives lenders sight of current-account inflows and outflows with a level of detail that M-Pesa SMS parsing only approximates. The analytical question is the same: what does income regularity, counterparty diversity, and spending pattern stability predict about repayment? The legal pipeline for getting that data is entirely different.

Second, consent architecture is not a checkbox exercise. Tala's app requests specific permissions and explains them in the local language at a fifth-grade reading level. In a GDPR environment, your lawful basis for processing alternative data signals must be documented, purpose-limited, and defensible to a Data Protection Authority. The consent moment is also where you lose borrowers who are uncomfortable with the scope of access you are asking for, which itself becomes a selection signal worth analyzing.

Third, the model governance requirements in Western markets are materially stricter. The EU AI Act's credit-scoring provisions, the FCA's consumer duty obligations in the UK, and the CFPB's adverse action notice rules in the US all require that a declined applicant receive an explanation they can actually act on. Building explainability into a model trained on 10,000 behavioral features is a genuine technical problem, not a compliance footnote.

Fourth, the data flywheel advantage is real but takes time to build. A lender entering thin-file underwriting in 2026 with no existing repayment data will have worse models than one with four years of labeled outcomes. Buying alternative data from a bureau aggregator (FinScore in Southeast Asia, Aire or Nova Credit in Western markets) lets you start faster, but you are renting someone else's signal rather than building your own.

The concrete lesson is this: Tala's advantage is not the algorithm; it is the willingness to define creditworthiness differently and then collect the data to support that definition. Data leaders who want to replicate the result need to start with the same question Siroya started with, which is not "what data do we have?" but "what behavior actually predicts repayment for this population, and how do we get lawful access to it?" The gap between those two questions is where most thin-file initiatives stall.

The full course on this sector:Data in Fintech.

Frequently asked questions

What alternative data does Tala use to underwrite loans?

Tala uses on-device smartphone signals collected with explicit user consent: M-Pesa SMS transaction records, call logs, app usage patterns and contact metadata such as frequency and reciprocity of communication. Its data scientists found the ratio of incoming to outgoing M-Pesa transactions and the diversity of counterparties predictive of repayment, variables that appear in no credit bureau file.

Are Tala's default rates high compared with traditional lenders?

Tala has cited default rates of 7 to 10 percent for its emerging-market portfolios, above the sub-3 percent a well-collateralized prime lender would accept. The relevant comparison is the alternative available to these borrowers: no formal credit at all, or an informal moneylender charging around 200 percent annualized.

Can a European lender copy Tala's model under GDPR and PSD2?

The principle transfers, the signals do not. Open banking access under PSD2 gives European lenders current-account inflows and outflows in more detail than M-Pesa SMS parsing, but the lawful basis for processing alternative data must be documented, purpose-limited and defensible to a Data Protection Authority, and the EU AI Act's credit-scoring provisions demand explainable declines.

What is the real competitive moat in alternative-data underwriting?

The data flywheel, not the algorithm. Repayment behavior from earlier cohorts feeds back into Tala's scoring model, so a lender starting thin-file underwriting in 2026 with no labeled outcomes will have weaker models than one with four years of history. Buying signals from aggregators like Nova Credit or FinScore speeds up launch but rents someone else's data.

Go deeper

The lessons that take this article further, free to read.

  1. 1Underwriting the thin-file customer with alternative dataData in fintech
  2. 2Consumer protection law: fair lending and disclosure rulesFintech: how the sector works
  3. 3Reading the transaction ledger: what payment and behavioral data revealData in fintech
  4. 4Data privacy and open finance: GDPR, CCPA and data-sharing consentFintech: how the sector works
  5. 5Bias, explainability & model cardsAI & machine learning strategy

Finished reading?

Validate your read to earn XP and feed your radar.