Tata Power Ajmer Distribution Limited  ·  Rajasthan, India  ·  2025
consumers scored
0
across three weighted detection layers

Predictive Meter Data Analytics for AT&C Loss Identification and Revenue Optimization

TPADL Theft Detection Engine v3.0 — Rule-based, interpretable, audit-ready

49,981HIGH risk consumers
₹6.99 CrBase annual saving
2.91×Benefit-cost ratio
+32.6%TPADL EBITDA impact
01 - Problem Statement

India's AT&C Loss Crisis and the TPADL Context

Aggregate Technical and Commercial (AT&C) losses represent the gap between energy a utility purchases and revenue it actually collects. India's national AT&C figure was 22.32% in FY 2020-21, appeared to decline to 15.41% by FY 2022-23, but reversed - the Power Finance Corporation's report for 2023-24 recorded a climb back to 16.12%. Accumulated sector losses stood at nearly ₹6,479 billion as of March 2023, an 11% increase year-over-year.

Commercial losses - the harder kind - are substantially driven by electricity theft. India loses an estimated $17 billion annually to non-technical losses (NTL), accounting for 30-40% of total generation. The global IEA average for comparable utility-level losses sits at around 7%. India is not a marginal outlier.

Traditional detection depends on reactive, field-based approaches: physical inspections triggered by complaints or random audits. There is no systematic mechanism to rank the full consumer base by theft probability. False-positive rates from complaint-driven inspections are high. The approach is expensive per confirmed case, and the fraction of actual theft caught is low. TPADL's estimated AT&C stands at approximately 18% - 3 percentage points above the RDSS national target of 15%.

The rationale of this study is to design a multi-layer, rule-based predictive detection framework that uses 12 months of smart meter data and MDI readings to identify consumers with high theft probability and generate a prioritized inspection list - entirely interpretable, field-auditable, and deployable without data science infrastructure.

$17 Bn
India electricity theft annual cost
30–40% of generation unaccounted; Kawoosa et al. (2023)
16.12%
National AT&C loss FY 2023-24
Reversed from 15.41% in FY 2022-23; PFC Report 2024
12–15%
RDSS national target
Revamped Distribution Sector Scheme benchmark for grants
18%
TPADL estimated AT&C baseline
3pp above RDSS target; this project targets the gap
02 - Dataset & Preprocessing

From 164,076 Records to a Clean Scoring Pool

12 months of monthly consumption (kWh) and MDI data, one record per consumer. Three rules run before any variable is computed.

164,076
Raw consumer base
P1
Null ≤10 kWh months
P2
Require ≥6 valid months
111,792
Scored pool (68.1%)
+
52,284
Excluded watchlist (31.9%)
P1 - Noise Exclusion
Monthly readings at or below 10 kWh are set to NULL. These reflect meter transmission failures, partial disconnections, or data entry placeholders - not real consumption. Treating them as zero would distort trend slopes and inflate volatility scores.
P2 - Minimum Valid Months
A consumer must have at least 6 valid months after P1 to enter the scoring pool. Below that, statistical calculations become unreliable. Excluded consumers route to an Insufficient Data watchlist and re-enter automatically as data accumulates.
P3 - Coverage Ratio
Computed for every consumer who passes P2:

Coverage_Ratio = Valid_Month_Count / 12

Feeds directly into the Zero Coverage Suspicion Score (ZCSS) in Layer 2.
Scored pool by rate category
DS-LT1
84.95%  94,970
NDS-LT2
13.04%  14,574
SP-LT5
1.20%  1,339
PSL-LT3
0.36%  399
AG-MS-LT4
0.33%  372
Others
0.12%  138
03 - Detection Architecture

Three-Layer Scoring Framework

Every consumer receives a Final Score from 0–100 by combining three weighted detection layers. All weights, thresholds, and formulas are explicitly documented and field-auditable.

Layer 1 — Consumption Pattern Index (CPI)
Consumption behaviour, seasonal patterns, volatility, peer comparison - 12 data points per consumer.
40%
CTS 20%
CVI 20%
DSR 15%
ZSB 10%
SCAS 15%
FCS 10%
PDS 10%
Layer 2 — Load Compliance Index (LCI)
Consistency between recorded consumption, sanctioned load, and Maximum Demand Indicator readings.
35%
LFR 30%
MOI 25%
MCGI 30%
ZCSS 15%
Layer 3 — Violation & History Score (VHS)
Prior enforcement history - most direct signal in the dataset. Step function, not linear.
25%
Violation Count step fn
Final Risk Score Final_Score  =  (0.40 × CPI)  +  (0.35 × LCI)  +  (0.25 × VHS)
All layer scores normalised 0–100 · Final_Score range: 0 to 100
Classification thresholds
HIGH
Score ≥ 65
7-day field inspection SLA. Financial leakage estimate generated.
MEDIUM
35 – 64
30–45 day scheduled inspection. Cross-check MDI vs load.
LOW
Score < 35
Quarterly watchlist. No immediate field action.
EXCLUDED
≤ 5 months
Insufficient data. Routed to meter read resolution team.
Override rules - escalate to HIGH regardless of score
Override O1 - Triggered: 0 consumers
Violation_Count ≥ 2
AND CPI ≥ 55
Zero fires because 99.5% of consumers carry no prior violations. Validates the case for behavioral detection over history-dependent rules.
Override O2 - Triggered: 49,859 consumers
MDI_Ratio ≥ 2.0
(peak demand ≥ 2× sanctioned load)
Drawing double contracted load is a hard compliance breach. Accounts for 99.8% of all HIGH classifications.
Override O3 - Triggered: 122 consumers
Sanctioned_Load ≥ 25 kW
AND Annual_LFR < 0.03
Large connection drawing near-zero load: either dormant (remove it) or bypassed (inspect it). Either outcome requires a visit.
04 - Layer 1 Variables

CPI - Seven Detection Variables

Version 3.0 expanded from four to seven variables. Three new signals fill gaps the original CPI could not catch.

CodeWeightVariableWhat it Catches
CTS20%Consumption Trend SlopeOLS slope on (month, kWh) pairs. A sustained negative slope across 8–12 months suggests a bypass was installed and the meter has been recording less ever since. Category thresholds: Residential 20 | Commercial 80 | Industrial 250 kWh/month.
CVI20%Consumption Volatility IndexCoefficient of Variation. Erratically jumping profiles - artificial spikes followed by near-zero months - are the fingerprint of partial bypass or on/off tampering. Score fires when CV ≥ 0.30.
DSR15%Drop Severity RatioGap between peak and trough valid months as % of peak. Catches cases CTS misses: V-shaped recoveries and non-linear declines. Fires when (Max – Min)/Max ≥ 40%.
ZSB10%Z-Score Statistical BreakSingle-month consumption cliff. Fires when minimum-Z month < –2.0 AND prior month was near-normal (Z ≥ –1.0). Targets the abrupt drop when a bypass is installed mid-year.
SCASnew15%Summer Consumption Anomaly ScoreApr–Jul is Rajasthan's peak demand season (AC, coolers, refrigeration). A consumer whose summer mean is materially below their off-season mean is showing the exact pattern field teams associate with summer bypass. Threshold: <0.50 ratio for Residential/Commercial; <0.60 for Industrial.
FCSnew10%Flat Consumption ScoreCorrects a blind spot: near-zero CoV was being treated as safe. A perfectly flat meter reporting identical readings every month on a large connection is one of the clearest signs of a fixed bypass or tampered register. Fires only when Mean > 50 kWh.
PDSnew10%Peer Deviation ScoreGroups consumers by Rate Category + sanctioned load bracket (nearest 5 kW). Uses LOW-risk consumer median as the clean baseline. Scores consumers >20% below peer median. The only variable with an external reference point - catches cases where all individual signals look moderate but absolute consumption is anomalous.
CPI Variable Weight Distribution
What this shows: CTS and CVI each carry 20% - the backbone of the index. The three v3.0 additions (SCAS + FCS + PDS) together account for 35%, directly filling the three blind spots the original 4-variable CPI had. Hover each segment to see exact weight.
Why the Weights Changed in v3.0
BLIND SPOT 1 → SCAS (15%)
No seasonal check - Rajasthan summer bypass invisible to CTS alone
BLIND SPOT 2 → FCS (10%)
Flat consumption scored as safe - fixed bypass register undetectable
BLIND SPOT 3 → PDS (10%)
No peer reference - absolute under-recording invisible without external benchmark
05 - Layer 2 Variables

LCI - Load Compliance Index

Four variables compare recorded consumption against what the connection's physical parameters imply. MDI (Maximum Demand Indicator) is the primary physical proxy.

CodeWeightVariableFormula & Signal
LFR30%Load Factor RatioAnnual_LFR = MEAN(kWh_m / (Sanctioned_kW × Hours_m)). Flags consumers systematically below category floor: Residential 6% | Commercial 8% | Industrial 10%. A 50 kW connection at 1% LFR is either dormant or bypassed.
MOI25%MDI Overdrawal IndexMDI_Ratio = Highest_MDI_kW / Sanctioned_Load_kW. Score fires above 1.05 (5% tolerance). Score=100 when MDI ≥ 2× sanctioned. A consumer whose consumption falls but MDI remains high is running load the meter is not recording.
MCGI30%MDI-Consumption Gap IndexMost innovative variable. Expected_kWh = MDI_m × Hours_m × LF_typical (0.20 Res / 0.35 Com / 0.50 Ind). MCGI = mean gap between expected and actual kWh. Score fires above 20% gap; reaches 100 at 80% gap. Closest available proxy for V-I mismatch without electrical parameter data.
ZCSS15%Zero Coverage Suspicion ScoreZCSS = (1 – Coverage_Ratio) × Load_Percentile × 100. Weights missing months by connection size. 8 missing months on a 1 kW household = low score. 8 missing months on a 50 kW industrial = high score. Separates genuine inactivity from suspicious gaps.
06 - Layer 3 Variables

VHS - Violation & History Score

One variable. A step function, not linear - the jump from zero to one violation is large. Beyond that, each additional violation adds progressively less because the consumer is already in the highest-risk tier.

Violation CountVHS ScoreInterpretation
00No recorded history. Score driven entirely by CPI and LCI.
140One prior flag. Enough to push a medium-CPI consumer into HIGH territory.
265Twice flagged. A known offender who has been penalised before.
382Three violations. Systematic repeat offender. Treat as HIGH regardless of consumption patterns.
≥ 4100Maximum VHS. Any consumer at this level goes to the top of the inspection queue.

Dataset context: 111,277 of 111,792 scored consumers (99.5%) carry zero prior violations. Only 515 consumers have any enforcement history (428 with one violation, 72 with two, 5 with three, 10 with four or more). Override O1 fired zero times because the joint condition - prior violations AND current CPI anomaly - requires a history that does not yet exist at scale. This finding validates the core argument: detection systems depending on historical violation records will systematically fail in Indian DISCOM contexts before systematic inspection coverage has been established.

Violation History Distribution - Why Override O1 Never Fired
Log scale required: The drop from 111,277 zero-violation consumers to just 428 with a single violation is so extreme that a linear scale makes all non-zero bars invisible. Only 87 consumers have 2+ violations - meaning Override O1's joint condition (violations ≥2 AND CPI ≥55) had almost no pool to draw from. Takeaway: any detection system built around historical enforcement will fail in the first inspection cycle. Behavioral scoring is the only viable approach.
07 - Detection Results

Risk Classification Across 111,792 Consumers

49,981HIGH Risk
44.71% of scored pool
550MEDIUM Risk
0.49% of scored pool
61,261LOW Risk
54.80% of scored pool
52,284EXCLUDED
Insufficient data watchlist
Risk Distribution (Scored Pool)
Primary Detection Signal - 111,792 Consumers
Key findings
Max Final Score: 59.77
No consumer reached the 65-point HIGH threshold through pure scoring alone. Every HIGH-classified consumer was elevated by an override rule (O2 or O3), not by computed CPI+LCI+VHS. This reflects data constraints - no billing data, no V/I readings - not a framework design flaw. With Phase 2 data, pure-score HIGH cases will emerge.
MOI Dominates at 40.6%
45,389 consumers have MDI Overdrawal as their primary signal - peak demand exceeds double their sanctioned load. Combined with DSR (37.6%), these two signals account for 78.2% of all classifications. The MDI signal is the most powerful available indicator in the current dataset.
Agriculture: 31% flagging rate
AG-MS-LT4 (agricultural connections) shows the highest average Final Score (25.2) of any rate category and a 31% HIGH flagging rate - highest proportion in the pool. SCAS and FCS are specifically tuned to separate legitimate seasonal inactivity from suspicious low-consumption patterns during active irrigation months.
HIGH Risk Flagging Rate by Rate Category
Proportional view: DS-LT1 dominates by absolute count (45,471 consumers) but its 47.9% flagging rate isn't the highest. AG-MS-LT4 agricultural connections flag at 30.9% despite being only 0.33% of the pool - SCAS catches their summer irrigation bypass pattern directly. SP-LT5 (Special Purpose Industrial) has the lowest rate at 6.8%, suggesting more stable consumption behaviour. Hover each bar to see absolute HIGH consumer counts.

TOP 10,000 priority consumers: combined 7.15 MU of estimated stolen energy and ₹5.25 Crore gross annual revenue loss. The highest-ranked consumer (Rank 1) is an MP-LT6 connection with 12 prior violations - estimated monthly theft of 1,607.8 kWh, gross annual loss ₹1,35,059, inspection ROI 2,601%.

08 - Financial Analysis

Revenue Leakage, AT&C Scenarios, and Investment Case

All tariff rates sourced from AVVNL Tariff for Supply of Electricity 2023. Financial module uses peer-benchmark and AT&C pool methodologies. Recovery factor: 60%.

₹3.56 Cr Gross annual revenue loss detected
(peer-benchmark, 7.77 MU stolen energy)
1,622 of 50,531 flagged consumers below peer median
₹2.13 Cr Recoverable after 60% recovery factor
(penalties, reconnection, write-off adjustment)
Conservative first-cycle estimate
15.26% Detection capture rate
(share of ₹23.29 Cr commercial loss pool identified)
Phase 2 with billing + DT data targets 65%+ capture
₹23.29 Cr Total commercial loss pool
(326.19 MU × 18% AT&C × ₹7.14/kWh blended tariff)
AVVNL 2023 tariff, weighted by consumer mix
₹7.14 Blended tariff (₹/kWh)
weighted average across detection dataset
AVVNL Tariff Order 2023 - actual rates
326.19 MU Estimated annual distribution volume
(TPADL service territory)
Base for AT&C scenario revenue recovery calculations
AT&C Reduction Scenarios
Scenario AT&C Target Annual Saving 5-Year NPV 10-Year NPV PAT Impact ARR Improve
Conservative16.5%₹3.49 Cr₹13.25 Cr₹21.47 Cr₹2.62 Cr₹0.11/unit
Base RDSS TARGET15.0%₹6.99 Cr₹26.49 Cr₹42.94 Cr₹5.24 Cr₹0.21/unit
Optimistic13.0%₹11.65 Cr₹44.15 Cr₹71.57 Cr₹8.74 Cr₹0.36/unit
Best Case12.0%₹13.98 Cr₹52.98 Cr₹85.88 Cr₹10.48 Cr₹0.43/unit
NPV at 10% WACC. PAT impact at 25% corporate tax rate. AT&C baseline: 18%. ARR improvement per unit distributed.
Annual Savings by Scenario (₹ Crore)
10-Year NPV by Scenario (₹ Crore)
Cumulative NPV Build-up - All Scenarios Over 10 Years (₹ Crore, 10% WACC)
Key insight: Every scenario turns NPV-positive in Year 1 - the programme pays back within the first inspection cycle. The gap between Conservative and Best Case widens each year; this is the compounding effect of sustained AT&C reduction. The shaded area under the Base scenario (RDSS target) shows how ₹6.99 Crore annual saving builds to ₹42.94 Crore in present-value terms by Year 10. Hover any point to compare all four scenarios simultaneously.
Programme investment case (Base Scenario)
₹24.02 Cr10-Year Net Programme NPV
2.91×Benefit-Cost Ratio
8.68 moApproximate Payback Period
₹15.16 CrFull Programme Cost (50,531 × ₹3,000)
Break-Even Thresholds
58.35 kWh/month - Year-1 cash recovery threshold. 3,373 of 50,531 flagged consumers exceed this (6.7% are Year-1 cost-positive).
3.50 kWh/month - Perpetuity basis threshold. Virtually every flagged consumer exceeds this. Theft stopped today eliminates loss in every future year. The inspection investment is one-time; the recovery is permanent.
TPADL Subsidiary EBITDA Impact
Sensitivity Analysis - Base Scenario (18% → 15%)
Annual Savings Sensitivity Matrix - AT&C Reduction × Tariff Rate (₹ Crore/year)
How to read: Each cell = annual saving from combining a specific AT&C reduction target (rows) with a tariff rate (columns). Darker blue = higher saving. The ★ row is the Base Scenario (18%→15% at ₹7.14 = ₹6.99 Crore). No cell is negative - the programme generates positive returns under every combination. Even the most conservative cell (1pp at ₹5.50/kWh) yields ₹1.79 Crore/year. Hover any cell for the exact figure.
Tata Power Group financial context (FY 2025)
Group Revenue
₹65,478 Cr
FY 2025 consolidated
Group EBITDA
₹15,354 Cr
23.45% EBITDA margin
Group PAT
₹3,971 Cr
6.06% net margin
TPADL EBITDA
+32.6%
Base scenario subsidiary improvement
09 - Conclusion & Contributions

Three Distinct Contributions to the Literature

The framework deviates from existing studies in three specific, documented ways - not incremental adjustments but architectural choices driven by field-operational constraints.

Contribution 01
Interpretable Rule-Based Architecture
Instead of a black-box ML model trained on a foreign benchmark dataset, the framework uses explicit, auditable rules that billing analysts and divisional managers can read, interpret, and calibrate without data science expertise. Every formula, weight, and threshold is documented with its operational rationale - a prerequisite for legal enforceability in Indian revenue tribunals.
Contribution 02
Network Validation via ZCSS and Override Rules
The framework integrates coverage-weighted zone analysis (ZCSS) and load-scaled override conditions as network-wide validation mechanisms. Override O3 (25 kW+ connection with <3% LFR) triggers mass survey protocols for large connections drawing near-zero load - an integration not formalized in prior single-study frameworks.
Contribution 03
Field-Linked Financial Quantification
The financial module directly links detection outputs to DISCOM-level metrics: ACS-ARR gap, EBITDA impact, PAT addback, inspection ROI, NPV, BCR, and consumer-level break-even thresholds. This end-to-end financial layer - absent from almost all academic detection models - converts an analytical output into a capital allocation decision framework.
Phase 2 future scope - data unlocks
Monthly Billed kWh + DT Mapping
Unlocks CDI, BMR, and transformer-level energy balancing. Detection capture rate estimated to rise from 15.26% to 65%+. Highest priority data request.
📊
Hourly / Daily Consumption Data
Enables peak-vs-offpeak ratio, weekday-weekend structural break, finer Z-score analysis. Required for temporal evasion pattern detection.
🔌
Voltage, Current & Power Factor Readings
Full electrical parameter layer: V-I mismatch, phase imbalance detection, power factor drop analysis. CT tampering detection not possible without these.
🤖
Hybrid ML Integration (Phase 3)
As confirmed-theft labels accumulate from field inspections, supervised XGBoost or LSTM can be trained on top of the rule-based scores. Rules remain interpretable; ML improves recall at the margin.

Programme result summary: 111,792 consumers scored → 49,981 HIGH (44.71%) → ₹6.99 Crore annual saving under base scenario → 32.6% TPADL EBITDA improvement → ₹24.02 Crore net 10-year NPV → 2.91× BCR → payback in 8.68 months. Every sensitivity scenario across all sixteen parameter combinations tested produces positive NPV.