This study examines the impact of digital financial literacy (DFL) on digital financial practices (DFP) in India by combining causal inference and machine learning approaches. It investigates whether higher DFL translates into greater engagement with digital financial services and whether this effect varies by gender.
Using primary survey data from 1,215 adults, the study applies covariate-adjusted ordinary least squares (OLS), propensity score matching (PSM) and overlap weighting to estimate the effect of DFL on DFP. Gender heterogeneity is analysed through interaction models and stratified matching. A Random Forest model is used to assess predictive importance.
Results consistently show that higher DFL is associated with significantly greater DFP across all estimators, with effect sizes ranging from 0.86 to 1.41 (p < 0.001). The effect is stronger among women, indicating a potential equalising role of literacy in digital finance adoption. Machine learning results further identify DFL as the most important predictor of DFP, explaining a substantial share of variation in outcomes.
The cross-sectional design limits causal interpretation, and findings are specific to the Indian context. Future research can extend this work using longitudinal data and comparative frameworks.
The findings highlight the need to shift from infrastructure-led to capability-driven financial inclusion strategies. Targeted DFL interventions, particularly for women and low-income groups, can improve adoption, trust and responsible use of digital financial services.
This study integrates causal inference and machine learning to provide robust evidence on the role of DFL in shaping digital financial behaviour, offering both methodological and policy-relevant insights for advancing digital financial inclusion.
1. Introduction
In the past twenty years, the entire financial services sector has changed drastically due to the digital revolution, which has especially been observed in developing economies. With the advent of smartphones, access to the Internet and the development of Fintech in India, the use of DFS has increased manifold. However, it is noteworthy that this is just an advancement for the digitisation of the economy and not necessarily indicative of successful adoption and security of DFS. In fact, the question of access as opposed to usage has become a critical problem of financial inclusion.
In this context, digital financial literacy (DFL) becomes a key determinant factor. DFL entails certain multi-dimensional capacities that consist of financial skills, digital skills, risk perception and ability to make effective use of DFS. According to available literature, DFL plays a vital role in enabling people to benefit from access to DFS. Low levels of DFL increase the risks of fraud in digital financial transactions, abuse of digital credit systems and ineffective decisions. These risks have been highlighted in OECD’s (2023) report, which states that low levels of financial literacy lead to overspending and poor risk management in the digital arena.
In spite of an emerging academic corpus, several gaps remain outstanding. Firstly, most studies have employed correlation or cross-sectional methods that limit causal analysis in any way. People who are better educated or earn higher incomes, or have had more experience using digital technology, also demonstrate higher levels of DFL and greater use of digital banking tools, which hinders causal inference due to the possible presence of self-selection bias (Yadav and Banerji, 2024). Secondly, there is a tendency for scholars to consider DFL as a continuous variable while studying its impact and ignoring treatment effects related to discrete values. In particular, few scholars consider various categories of literacy, such as high vs low, to identify differences in treatment effects. Lastly, the effect of literacy on behaviour has rarely been explored from the perspective of heterogeneous impacts based on gender-specific access, confidence and usage differences (Yadav and Banerji, 2023).
These are areas that are particularly pertinent to the Indian situation. Although there have been significant advancements on the part of India in building digital infrastructure for finance, there still exist inequalities when it comes to the use of this technology based on socio-economic differences. Research carried out in Financial Literacy and Inclusion in India (2025) indicates the existence of inequalities where there is an imbalance in the connection between access and usage, particularly among poorer communities and rural regions. The problem of the gender gap is particularly evident. For instance, according to research conducted by Access and Usage of Financial Products in India: A Gender Gap Analysis (2023), women tend to use fewer digital financial products, regardless of their income, education levels and locations.
In light of these issues, the current article uses a multiple methods approach to investigate the causal relationship between DFL and DFP in India. This involves the use of covariate-adjusted ordinary least squares (OLS), propensity score matching (PSM) and overlap weighting methods in estimating the impact of DFL on DFP, thus enhancing the credibility of causal inference through methodological triangulation. Second, this analysis will consider the differential impacts of a high level of DFL (top third of the sample) and a low level of DFL (bottom third of the sample).
In addition to econometrics, this article adopts an approach that involves machine learning (ML) to aid in causal analysis. For instance, Random Forest models have been used to establish non-linear associations, measure the significance of variables and test the predictability of DFL on the behaviour of people’s finances. This method fills the gap in the literature by separating causal inference from predictions.
There are three key contributions made by this study. First, there is solid empirical evidence that DFL is an important variable in determining digital financial behaviour within one of the most extensive digital finance platforms in the world. Secondly, this study contributes to methodological practices through a novel combination of causal inference and ML approaches, providing a replicable approach to future research. Finally, this study provides policymakers with useful insights into the heterogeneous effects of DFL on different gender groups and literacy levels.
The findings have therefore helped to highlight the fact that financial inclusion in the digital age requires not just the enhancement of accessibility but also the capacity building of the users. The study has provided insight into this issue by integrating the issues of literacy, behaviour and heterogeneity in the analysis.
2. Literature review
2.1 Theoretical foundations of digital financial literacy and behaviour
The concept of DFL can be described as a combination of financial and digital literacy that emerged with respect to the challenges in a digitised financial environment (Yadav et al., 2025a, b). Based on the human capital theory (Becker, 1964), the idea of financial literacy suggests that education and skills improve individuals’ ability to make sound decisions and have better financial outcomes. Nonetheless, due to the ongoing digitalisation of the modern financial environment, the term has received a more extended interpretation. According to the OECD/INFE (2023) definition, DFL involves the safe and efficient use of digital financial services as well as awareness of possible cyber-related risks and data protection issues.
From the perspective of behaviourism, the theory of planned behaviour (Ajzen, 1991) is an excellent theoretical model that can explain the impact of literacy on financial behaviours. DFL contributes to improved behavioural control through building greater confidence in users in dealing with digital systems like mobile banking and fintech apps. Previous research (Mishra et al., 2024; Yadav and Banerji, 2025) indicates that digital literacy influences attitudes favourably and decreases perceptions of barriers to adoption. Yet, TPB research mostly focuses on the intention aspect rather than the actual behaviours.
The theory is further augmented by behavioural economics through the aspects of trust, risk perception and bounded rationality. When an individual embraces the use of digital finance, he/she is faced with the risks associated with the security and reliability of the platform. Trust becomes an important determinant in such instances (Gefen et al., 2003). Through DFL, individuals get the ability to identify the risks involved and identify reliable platforms. Though risk perceptions in relation to trust have been acknowledged in past literature, it is often considered as contextual rather than causative.
2.2 Determinants and measurement of digital financial literacy in India
In India, there is empirical research on DFL regarding its measurement and its socio-economic determinants. For example, in their study in 2022, Ravikumar et al. developed a multi-dimensional scale to measure DFL consisting of digital literacy, awareness of cyber risks, understanding consumer rights and digital responsibility. The findings revealed substantial variations across education level, income level and urban-rural dichotomy, indicating inequality in the distribution and determinants of DFL.
Pattnayak and Sahoo (2024) have found education, income, use of the Internet, gender and rural residence to be the important determinants of DFL. In the same vein, the study conducted by Azeez and Akhtar (2021) has emphasised the impact of education and occupation on adopting online banking services, especially in rural settings. All the above-cited studies corroborate the capability approach to understanding DFL, as well.
Although much progress has been made in terms of measuring DFL, most research has considered it to be an endogenous construct that is shaped by socio-demographic variables. While this approach provides useful insights, it fails to recognise the power of DFL as an exogenous factor that may shape financial behaviour.
2.3 Digital financial literacy and digital financial practices
There is an increasing number of studies that show the link between high DFL and adoption of digital financial transactions (DFTs), which include mobile payments, online banking and digital investments. According to Dube et al. DFL and awareness increase the likelihood of fintech adoption by Indian millennials. Likewise, Ravikumar et al. found that people who have high DFL are more active and selective users of digital financial systems.
Nevertheless, one major constraint of the research is its dependence on the application of correlation approaches. According to Azeez and Akhtar (2021), individuals who are more literate possess more confidence in using digital technologies, yet the effect of the reverse causation approach cannot be ignored. This implies that there has been an inadequate analysis of the relationship between DFL and DFP.
This gap is particularly significant in policy contexts, where establishing causality is essential for designing effective literacy interventions. The absence of robust identification strategies limits the ability to infer whether improving DFL directly translates into enhanced digital financial behaviour.
2.4 Emerging integration of machine learning and causal approaches
Innovations in the area of data analytics have led to the emergence of innovative ways of data analysis known as ML. Random forest algorithms and gradient boosting algorithms, among other techniques, have proved effective in dealing with the issue of nonlinearity and dimensionality. However, the application of ML in the area of financial literacy and behaviour is yet to be fully explored.
According to the bibliometric study conducted by Pattnaik et al. in an article that appeared in 2024, most of the applications of ML in finance were on areas such as stock prediction and fraud detections, but none had any relation to literacy and behavioural issues. One recent literature concerning financial literacy was conducted by Lu et al. (2024).
The limitation of the current body of knowledge lies in its fragmented nature, where some research adopts one form of approach while ignoring another. While the causality methods help to provide directionality, ML methods strengthen the predictability of results in financial analyses. It is necessary to fill this methodological gap for better insights into digital finance behaviour.
2.5 Gendered dimensions and heterogeneity in DFL–DFP relationship
Gender differences continue to be a source of worry when it comes to digital financial inclusion. While programs like the Jan Dhan Yojana have been successful in increasing digital financial inclusion, issues related to lack of literacy, self-confidence and utilisation still exist. Pattnayak and Sahoo (2024) noted that women show much lower DFI levels as a result of education and socio-cultural restrictions.
According to the Grant Thornton Bharat (2024) report, despite 86% of women in India having a bank account, only about half utilise digital finance. According to the Economic Survey (2023–24), there exist gender inequalities in usage intensity as opposed to accessibility, implying that literacy and not infrastructure constitutes the main barrier.
Nonetheless, the vast majority of empirical works consider gender as a control factor but not an explanatory one. Thus, there is little information on how DFL affects different genders. The question about the possible effect of gender heterogeneity on the impact of DFL is largely understudied.
2.6 Positioning of the present study
In bridging these gaps, this study contributes to the existing body of knowledge on DFL in the following manner. Firstly, it enhances the current research by proving a causal effect from DFL to DFP based on a number of different approaches, including covariate-adjusted OLS estimation, PSM and overlap weighting. Secondly, this research proposes the reconceptualisation of DFL as a predictor instead of a dependent variable.
Thirdly, this research includes gender heterogeneity in its analysis through the use of an interaction model and matching techniques. Fourthly, it bridges the gap between ML and causal inference.
Theoretical insights from the human capital perspective, TPB and behavioural economics, combined with sophisticated empirical approaches, are significant contributions of this research study. This research has provided convincing findings for policymakers and practitioners interested in developing capability-based digital financial inclusion initiatives.
2.6.1 Objectives
To examine the relationship between DFL and digital financial practices (DFP) in India using multiple econometric identification strategies, including covariate-adjusted OLS regression, PSM and overlap weighting, to generate credible causal evidence.
To estimate the causal impact of higher DFL (top tertile) relative to lower DFL (bottom tertile) on DFP, while controlling for potential confounding factors through matching and weighting techniques.
To investigate heterogeneity in the DFL–DFP relationship by gender by incorporating interaction terms in OLS models and conducting stratified PSM analyses, to assess whether improvements in DFL yield differential effects on men’s and women’s digital financial behaviours.
3. Research methodology
This study adopts a cross-sectional quantitative and explanatory research design rooted in the positivist paradigm to examine the relationship between DFL and DFP in the Indian context. The primary data are gathered using a structured online survey conducted to collect data using a stratified purposeful sampling method to ensure representation of demographics along the lines of gender, age group, income group and geographic location. After deleting the data points for missing and inconsistent responses and outliers, 1,215 were selected for analysis. Data screening involved the removal of incomplete responses, inconsistent entries and statistical outliers to improve data quality and reliability of estimation. DFL was measured using a multi-dimensional scale covering knowledge, skills, awareness and confidence in the use of digital financial tools. DFP was operationalised as a behavioural index comprising digital payments, online banking usage, fintech adoption and secure financial behaviour. All items were measured using a five-point Likert scale and were validated using reliability assessment through Cronbach’s alpha and composite reliability and construct validity through AVE. Reliability and validity statistics were within acceptable thresholds, supporting the consistency and construct adequacy of the measures used. To estimate treatment-effect approximations while reducing observable selection bias, the study employs a multi-method analytical framework comprising covariate-adjusted OLS regression, PSM and overlap weighting estimators. The application of PSM and overlap weighting relies on the assumptions of conditional independence (unconfoundedness) and common support. These assumptions were addressed by including relevant demographic, socio-economic and financial covariates in treatment assignment models. Post-estimation balance diagnostics indicated satisfactory comparability between treated and control groups. Nevertheless, as with all observational studies, potential bias arising from unobserved factors cannot be fully eliminated. Heterogeneity in estimated effects was examined through gender interaction models and stratified matching analyses. The Random Forest regression model was used to validate findings through predictive accuracy and variable-importance diagnostics. Ethical integrity was guaranteed by means of voluntary participation, anonymity and institutional consent procedures.
While these methods strengthen internal validity, the cross-sectional nature of the data limits definitive causal interpretation. Accordingly, findings should be interpreted as robust associative evidence with treatment-effect approximations rather than conclusive causality.
4. Data analysis
The descriptive statistics of both DFL and DFP are shown in Table 1. In relation to the mean score, the DFL has an average score of 3.307, with a standard deviation of 0.770. As for the DFP, its mean score is 3.189, with a standard deviation of 0.871, which is greater than half of the scale. This implies that the respondents have a high degree of positivity towards DFL and practice. Variation in both DFL and DFP shows that respondents vary in relation to their knowledge and familiarity with digital financial tools.
Descriptive statistics and correlation between DFL and DFP
| Variable | Minimum | Maximum | Mean | Median | SD | 1st quartile | 3rd quartile | Correlation (r) |
|---|---|---|---|---|---|---|---|---|
| DFL | 1.000 | 4.625 | 3.307 | 3.500 | 0.770 | 2.844 | 3.875 | 0.849 |
| DFP | 1.452 | 4.821 | 3.189 | 3.333 | 0.871 | 2.440 | 3.833 | – |
| Variable | Minimum | Maximum | Mean | Median | SD | 1st quartile | 3rd quartile | Correlation (r) |
|---|---|---|---|---|---|---|---|---|
| DFL | 1.000 | 4.625 | 3.307 | 3.500 | 0.770 | 2.844 | 3.875 | 0.849 |
| DFP | 1.452 | 4.821 | 3.189 | 3.333 | 0.871 | 2.440 | 3.833 | – |
The DFL range of 1.000–4.625 and DFP of 1.452–4.821 indicate that the responses vary across the participants, which can be linked to the fact that digital adoption and literacy levels are not evenly distributed. In fact, such dispersion is usually common in an emerging economy like India, where the differences in demographic, technological and socioeconomic aspects distort the pace of digital financial inclusion.
The strong positive relationship between DFL and DFP is demonstrated through the statistically significant value of the Pearson correlation coefficient, r = 0.849. It should, therefore, be expected that a higher level of DFL would result in increased levels of DFP such as online banking, mobile payments and activities of digital savings and investment. The strength of the relation is consistent with the previous literature that has noted that increased digital literacy would enhance individuals’ confidence and participation in the digital financial ecosystem (Koskelainen et al., 2023; Hasan et al., 2023).
The strong correlation further hints at the potential for DFL being one of the most important facilitators of behavioural change regarding financial matters. DFL will ensure that users comprehend digital security procedures better, shop around for the best financial products and services, and optimise their decision-making processes in the digital space (OECD, 2022). However, even though the strength of the correlation is evident, the lack of clear causation highlights the need to apply econometric models and techniques of causal inference, including PSM and overlap weighting, in order to discover the actual cause-and-effect relationship between DFL and DFP.
Furthermore, the findings lend more weight to the theoretical assumptions underlying the Capability Approach (Sen, 1999) and technology acceptance model (TAM) (Davis, 1989), which state that knowledge and perceived ability impact the behavioural intention and actions of an individual. In light of the above-mentioned, it can be argued that DFL serves as one of the main capacity builders, as it helps people take informed and competent action when making financial decisions. The DFL-DFP relationship emphasises the significance of digital capability-building efforts.
The observed statistical characteristics provide a basis for further causal inference. Due to the large degree of variability in literacy levels and the clear relationship, this set of data is very well suited to estimation via OLS regression, PSM and overlap weighting models, and so on. Subsequent tests for heterogeneity for gender and other categories would aid in determining whether such a relationship differs between various groups, which would prove to be pertinent in the context of the digital divide existing in India (Kulkarni and Ghosh, 2021).
Table 2 presents the range of causal estimates for the impact of DFL on DFP, derived using three complementary identification strategies: baseline OLS regression model, PSM and Overlap Weighting. First, the robust and positive statistical significance of results across the different models suggests that there is a strong and consistent positive relationship between the levels of literacy and the practices related to digital finance engagement.
Estimated effects of digital financial literacy on digital financial practices
| Model | Effect estimate | Robust SE | 95% CI (L–U) | Significance |
|---|---|---|---|---|
| Baseline OLS | 0.857 | 0.056 | 0.747–0.968 | ***p < 0.001 |
| PSM ATT (exact matching) | 1.305 | 0.182 | 0.948–1.662 | ***p < 0.001 |
| Overlap weighting (ATO) | 1.412 | 0.123 | 1.171–1.653 | ***p < 0.001 |
| Model | Effect estimate | Robust SE | 95% CI (L–U) | Significance |
|---|---|---|---|---|
| Baseline OLS | 0.857 | 0.056 | 0.747–0.968 | ***p < 0.001 |
| PSM ATT (exact matching) | 1.305 | 0.182 | 0.948–1.662 | ***p < 0.001 |
| Overlap weighting (ATO) | 1.412 | 0.123 | 1.171–1.653 | ***p < 0.001 |
Note(s): OLSs = Ordinary least squares; PSM = Propensity score matching (average treatment effect on the treated, ATT); ATO = Average treatment effect for the overlap population. Robust standard errors were used to control for heteroscedasticity
4.1 Baseline OLS estimates
The effect estimate from the baseline OLS regression is 0.857 (p < 0.001), indicating that for every one-unit increase in DFL, there is an approximate 0.86-point increase in DFP, assuming all other covariates are held constant. This strong effect aligns with existing literature, which indicates that a higher level of digital literacy improves financial management, enhances transaction efficiency and fosters trust in digital systems (Ravikumar et al., 2022). A narrow 95% CI ranging from 0.747 to 0.968 reflects the precision and reliability of this estimate, thus reinforcing the hypothesis that DFL significantly predicts behavioural outcomes in digital finance.
However, while informative, the OLS estimates may suffer from possible selection bias and endogeneity – persons with higher financial awareness may self-select into the usage of digital financial platforms. Hence, to establish more credible causality, advanced quasi-experimental methods have been applied.
4.2 Propensity score matching (PSM) results
The PSM analysis, undertaken within the ATT framework, estimates the treatment effect to be 1.305 (p < 0.001). That is, after accounting for the relevant observable covariates-age, gender, income, education and sector of employment-individuals in the top tertile of DFL have average DFP scores 1.31 points higher compared to those in the bottom tertile. The result illustrates that when matched on comparable socio-economic characteristics, individuals with higher literacy still perform significantly better in digital financial behaviours.
This result gives significant importance to the role of DFL in enhancing financial engagement and inclusion, especially in the context of the rapidly digitalising financial ecosystem of India. It also underlines the enabling capability notion of DFL, that is, a concept central to Sen’s Capability Approach (1999), which purports that knowledge and competence create new possibilities for individuals to make informed and independent choices. An increase in the effect magnitude, as derived through PSM (1.305) from the one derived using OLS (0.857), signifies that this causal estimate turns out to be higher when selection bias is accounted for after balancing the covariates.
Figure 1 shows the mean differences (SMDs) of all covariates that were controlled before and after propensity score calliper matching. Before matching, some covariates, especially age and education levels, and income groups, there was a significant imbalance in high and low DFL groups, meaning the SMD values were above the generally accepted value of 0.1. After PSM using callipers, standardised mean differences of all covariates were below 0.1, which means there was excellent post-matching balance of treatment and control groups. This implies that the respective procedure was effective in forming similar groups in terms of perceived demographic and socio-economic traits. Therefore, the estimates are less prone to observable confounding and may be interpreted as stronger associative evidence consistent with treatment effects.
The x-axis measures absolute standardized mean difference from 0.0 to 1.2, and the y-axis lists covariates. Triangles represent values before matching, and circles represent values after matching. A dashed line at 0.1 marks the threshold for acceptable balance. After matching, all covariates fall below this threshold, indicating improved balance.Post-matching covariate balance diagnostics indicate substantial improvement in standardised mean differences across treatment groups. Source: Authors’ own work
The x-axis measures absolute standardized mean difference from 0.0 to 1.2, and the y-axis lists covariates. Triangles represent values before matching, and circles represent values after matching. A dashed line at 0.1 marks the threshold for acceptable balance. After matching, all covariates fall below this threshold, indicating improved balance.Post-matching covariate balance diagnostics indicate substantial improvement in standardised mean differences across treatment groups. Source: Authors’ own work
Figure 2 shows covariate balance obtained using PSM using exact matching constraints. High imbalances were present across a range of covariates before the adjustment, especially age, education level and income category, as indicated by a large standardised mean difference between the high and low DFL groupings. Following the use of exact matching together with propensity score adjustment, standardised mean differences of all covariates were narrowed down to values that were very near zero and far less than 0.1. It implies almost complete covariate balance and shows that there is no difference in treatment and control groups except that these groups are virtually identical regarding observed demographic and socio-economic characteristics. This high balance strengthens confidence that observed covariate differences are substantially reduced, improving the credibility of treatment-effect estimates.
The scatterplot measures absolute standardised mean differences across covariates. The x-axis represents the absolute standardised mean difference, ranging from 0.0 to 1.4. The y-axis lists various covariates such as distance, age_num, education_level_3, and others. Unmatched data points are shown as squares, and matched data points are shown as diamonds. Most matched data points are near zero, indicating substantial improvement in covariate balance after exact matching.Exact matching substantially improves covariate balance, reducing standardised mean differences across treatment groups. Source: Authors’ own work
The scatterplot measures absolute standardised mean differences across covariates. The x-axis represents the absolute standardised mean difference, ranging from 0.0 to 1.4. The y-axis lists various covariates such as distance, age_num, education_level_3, and others. Unmatched data points are shown as squares, and matched data points are shown as diamonds. Most matched data points are near zero, indicating substantial improvement in covariate balance after exact matching.Exact matching substantially improves covariate balance, reducing standardised mean differences across treatment groups. Source: Authors’ own work
Figure 3 shows the covariate balance diagnostics calculated with the help of an overlap weighting (average treatment effect for the overlap population, ATO) by propensity scores. Some of the covariates, especially age, education levels and income groups, before adjustment, had high and low groups of DFL that were highly imbalanced, as indicated by large, standardised differences. The overlap weighted mean standardised difference of all the covariates was brought to values that are nearer to zero and far less than the traditional set point of 0.1, which is a way of expressing superb balance. In contrast to matching methods, which ignore observations, overlap weighting gives more weight to those with similar propensity scores and concentrates inference on the areas of overlap between treatment and control groups. A high post-weighting balance vindicates that the overlap weighting method is effective in reducing evident confounding and yielding more credible treatment-effect approximations under stated assumptions of DFL on DFP.
The x-axis measures absolute standardized mean difference from 0.0 to 1.2, and the y-axis lists covariates. The graph compares unweighted and overlap weighted samples. The unweighted sample shows larger differences, while the overlap weighted sample shows differences closer to zero. The propensity score is at zero, and age_num is around 0.1. Education levels and income groups show varying degrees of balance improvement. All values are approximated.Overlap weighting substantially improves covariate balance, with post-weighting standardised mean differences approaching zero across covariates. Source: Authors’ own work
The x-axis measures absolute standardized mean difference from 0.0 to 1.2, and the y-axis lists covariates. The graph compares unweighted and overlap weighted samples. The unweighted sample shows larger differences, while the overlap weighted sample shows differences closer to zero. The propensity score is at zero, and age_num is around 0.1. Education levels and income groups show varying degrees of balance improvement. All values are approximated.Overlap weighting substantially improves covariate balance, with post-weighting standardised mean differences approaching zero across covariates. Source: Authors’ own work
4.3 Overlap weighting (ATO) estimates
The overlap weighting (ATO) (Table 2 and Figure 3) method gives an even larger estimate of 1.412 (p < 0.001), with a 95% CI ranging from 1.171 to 1.653. This reweights the observations to focus on people in areas of overlap in covariates, enhancing comparability and generalisability (Li et al., 2018). The size of this effect suggests that, for people with similar backgrounds and moderate likelihood of treatment, the higher the DFL, the more it improves DFP by roughly 1.4 points.
The robustness of this finding across these econometric models presents robust empirical evidence supportive of the first two objectives of this study. This clearly establishes that higher DFL alone contributes directly to improved practices for the same, and not just as an outcome of correlation but as an important explanatory mechanism associated with behavioural improvement. The statistical consistency across OLS, PSM and ATO addresses the concerns of model dependence and confounding, therefore giving strong validity to the inference.
4.4 Gender-based heterogeneity and next steps
While Table 2 presents the overall causal effect, further analyses, informed by Objective 3, will evaluate whether these effects are significantly different between men and women. Previous literature suggests that digital gender divides continue to exist in India, influenced by cultural norms, technology access and confidence (Kulkarni and Ghosh, 2021; Arora, 2021). Thus, the interaction effects of DFL*Gender and stratified PSM analysis will provide insight into whether interventions need to be differentiated to address such inequities in digital financial behaviour.
Table 3 reports the results of gender-stratified PSM, which examined whether the estimated effect of DFL on DFP does indeed vary across men and women. Both coefficients are positive and highly significant (p < 0.001), confirming that higher levels of DFL significantly enhance DFP in both groups. However, the magnitude of the effect varies notably between the two groups, indicating meaningful gender-based heterogeneity in the DFL–DFP relationship.
Gender-based heterogeneity in the estimated effect
| Gender | ATT (high vs low DFL) | Std. Error | t-value | p-value |
|---|---|---|---|---|
| 1 (Male/Reference) | 1.31 | 0.219 | 5.98 | <0.001 |
| 2 (Female) | 1.61 | 0.326 | 4.93 | <0.001 |
| Gender | ATT (high vs low DFL) | Std. Error | t-value | p-value |
|---|---|---|---|---|
| 1 (Male/Reference) | 1.31 | 0.219 | 5.98 | <0.001 |
| 2 (Female) | 1.61 | 0.326 | 4.93 | <0.001 |
Note(s): ATT = Average treatment effect on the treated. Estimates were obtained using stratified propensity score matching (PSM) to compare individuals in the top and bottom tertiles of digital financial literacy (DFL) within gender subgroups. Robust standard errors are reported
4.5 Male participants
For male respondents, the estimated ATT value is 1.31 (t = 5.98, p < 0.001). This indicates that, on average, men in the top tertile of DFL have a digital financial practice level 1.31 points higher than those in the bottom tertile, given observables like age, income, education and employment sector.
The implication is that the average male user shows more engagement in digital financial platforms, reflecting a higher level of confidence, exposure and experience with technology-based financial systems Chhillar and Arora (2022). Men are also more likely to have access to smartphones, financial tools or digital means of payment. This is because it again increases the probability of translating literacy skills into practice. But a considerable treatment effect is also stressing the point that digital literacy is a vital determinant, even in the context of males, which again asserts additional benefits regardless of gender.
4.6 Female participants
For females, the ATT is 1.61 with t = 4.93 and p < 0.001, indicating it is greater in magnitude compared to males. This shows there is an even bigger advantage for females in having better DFL to enhance their digital financial behaviours. Another interpretation is that the difference in DFP will be bigger for females compared to males because the increase in the level of DFL will lead to an increase in the level of DFP.
From a policy perspective, it is highly pertinent to digital financial inclusion from a gender equality angle. It can be assumed that literacy efforts tailored for women could lead to higher behaviour-focused outcomes compared to efforts for men. This corresponds with global trends regarding digital literacy, filling in both financial and social gaps for women to join decision-making processes, savings and entrepreneurship through digital channels (OECD, 2022).
4.7 Comparative and causal insights
The comparison across male and female groups shows that while both have statistically significant elasticities, the elasticity of DFL for females is higher than that of males: 1.61 against 1.31. This points to the fact that DFL is a more powerful catalyst for behaviour modification among women, partly because they start from a relatively lower baseline of access or confidence. Therefore, the marginal benefit of improvement in literacy is greater among female participants.
The observed heterogeneity thus provides an additional justification for the interaction model results, in which the DFL × Gender term is expected to bear a positive and significant coefficient in regression frameworks, further confirming that the behavioural effect of DFL is not uniform across genders. This confirms Objective 3 of the study, which sought to evaluate whether the improvements in DFL have different effects across demographic groups.
The variable importance findings of the Random Forest model to predict the DFP are shown in Figure 4. It should be noted that DFL (DFL_pure) is the most significant predictor of DFP, with the highest percentage change in mean squared error (%IncMSE ≈ 47.6) and change in node purity. What this means is that eliminating or swapping DFL causes the most significant loss in the predictive capability of the model. The next most significant predictors are education level and age, which means that the human capital and life-cycle factors can also be seen as influencing the development of DFP. The predictive importance of income level and gender is relatively less, which indicates that after accounting for literacy, education and age, their contribution to the prediction of DFP is not very substantial. Overall, the findings of the Random Forest support the econometric results by revealing DFL as the most powerful predictor of digital financial behaviours, as well as providing the relative priority of socio-demographic covariates in a non-parametric, data-driven model.
The x-axis lists variables and the y-axis measures the percent increase in mean squared error. DFL pure has the highest importance at 47.6 percent, followed by education level at 28.2 percent, age number at 26.2 percent, income level per month at 19.3 percent, and gender at 7.5 percent.Random Forest variable importance identifies digital financial literacy as the strongest predictor of digital financial practices, followed by education and age. Source: Authors’ own work
The x-axis lists variables and the y-axis measures the percent increase in mean squared error. DFL pure has the highest importance at 47.6 percent, followed by education level at 28.2 percent, age number at 26.2 percent, income level per month at 19.3 percent, and gender at 7.5 percent.Random Forest variable importance identifies digital financial literacy as the strongest predictor of digital financial practices, followed by education and age. Source: Authors’ own work
Figure 5 illustrates the partial dependence of DFL on predicted DFP. The relationship is positive and non-linear, indicating that increases in literacy are associated with higher engagement in digital financial behaviour. The sharper rise at moderate literacy levels suggests threshold effects, while the plateau at higher levels indicates diminishing marginal gains.
The x-axis measures digital financial literacy from 2 to 4, and the y-axis measures predicted digital financial practices from 2.75 to 3.75. The data line shows a positive non-linear relationship, with a sharper rise at moderate literacy levels and a plateau at higher levels.Partial dependence plot from the Random Forest model showing a positive non-linear relationship between digital financial literacy and predicted Digital Financial Practice. Source: Authors’ own work
The x-axis measures digital financial literacy from 2 to 4, and the y-axis measures predicted digital financial practices from 2.75 to 3.75. The data line shows a positive non-linear relationship, with a sharper rise at moderate literacy levels and a plateau at higher levels.Partial dependence plot from the Random Forest model showing a positive non-linear relationship between digital financial literacy and predicted Digital Financial Practice. Source: Authors’ own work
Table 4: The Random Forest Regression technique was utilised for validating the econometric analysis in respect of any possible non-linear relationship between DFL and DFP. The high R2 score of 0.687 highlights the strength of this technique regarding variance explained; in other words, it can be assumed that about 69% of the variance in the DFP of the respondents is accounted for by the set of independent variables. In addition, the root mean square error (RMSE) score of 0.523 indicates that there is a high level of predictive accuracy with minimal errors.
Random forest regression results for predicting digital financial practices (DFP)
| Metric/Variable | Description | Result | Interpretation |
|---|---|---|---|
| Model type | Random Forest regression (ntree = 1,000) | – | Non-parametric ML approach for predictive validation |
| R2 (test set) | Coefficient of determination | 0.687 | Model explains ∼68.7% of the variance in DFP |
| RMSE (test set) | Root mean square error | 0.523 | Indicates good predictive accuracy |
| DFL (%IncMSE) | Increase in mean squared error when variable is permuted | 47.6% | DFL = most important predictor |
| Education level (%IncMSE) | 28.2% | 2nd strongest predictor | |
| Age (%IncMSE) | 26.2% | 3rd strongest predictor | |
| Income level (%IncMSE) | 19.3% | Moderate influence | |
| Gender (%IncMSE) | 7.5% | Least influence | |
| Partial dependence trend | DFL → Predicted DFP | Monotonic Positive | Confirms consistent behavioural improvement with higher literacy |
| Metric/Variable | Description | Result | Interpretation |
|---|---|---|---|
| Model type | Random Forest regression (ntree = 1,000) | – | Non-parametric ML approach for predictive validation |
| R2 (test set) | Coefficient of determination | 0.687 | Model explains ∼68.7% of the variance in DFP |
| RMSE (test set) | Root mean square error | 0.523 | Indicates good predictive accuracy |
| DFL (%IncMSE) | Increase in mean squared error when variable is permuted | 47.6% | DFL = most important predictor |
| Education level (%IncMSE) | 28.2% | 2nd strongest predictor | |
| Age (%IncMSE) | 26.2% | 3rd strongest predictor | |
| Income level (%IncMSE) | 19.3% | Moderate influence | |
| Gender (%IncMSE) | 7.5% | Least influence | |
| Partial dependence trend | DFL → Predicted DFP | Monotonic Positive | Confirms consistent behavioural improvement with higher literacy |
Among all the predictor variables, the %IncMSE for DFL was the highest at 47.6% when permuted, implying that DFL strongly drives digital financial participation. This finding is supported by the econometric result that the capability of individuals to take up and use digital financial services and place their trust in these platforms is a critical factor in their digital finance literacy. It can also be observed from the partial dependence plot that the relationship between DFL and DFP is monotonic and positive.
Although the level of education emerged as the second major factor, with %IncMSE = 28.2%, previous studies have highlighted the need to emphasise the role of education in enhancing an individual’s technological skills while making sound financial decisions (Sharma and Kumar, 2025). The third major factor that emerged was the age variable, %IncMSE = 26.2%, with evidence of a learning curve across generations, where the older, less technologically savvy generation appears to be increasing in adaptability due to exposure.
Income level also had a significant yet moderate effect (%IncMSE = 19.3%), which implies that the higher the financial ability and the availability of digital devices, the higher the likelihood of increased participation within the digital finance system. On the other hand, gender had the lowest predictive capability (%IncMSE = 7.5%), indicating that despite the persistence of the gap between the two genders, the gender gap within the digital finance system is gradually narrowing due to the focus of inclusive policies and selective literacy schemes (Yadav et al., 2025).
From the methodologies involved, it is evident that the ML results validate the robustness of the estimates of causality obtained from the conventional econometric methods, such as OLS analysis, analysis involving PSM and overlap weighting. This is attributed to the non-parametric nature of the Random Forest models, which handles all the linear assumptions and multicollinearity. Hence, the result obtained from either ML or conventional econometric methodologies establishes a robust conclusion regarding the positively significant influence of DFL on the DFP.
5. Discussion
Digital finance systems and practices that have developed in the developing nations of the world, like the Indian economy, have proved to create a significant influence on the way people engage with finance, savings and payment systems. In this field of Digital Finance Practice, an important role is performed by DFL, which functions as the catalytic element in Digital Finance Practices that makes the realms of inclusion and resilience attainable. This research makes a contribution to the still emerging body of evidence regarding the presence of a cause-and-effect relationship between DFL and Digital Finance Practices in a complementary manner through the fusion of both Econometric and Machine Learning Models with a view to overcoming the concerns of robustness. This research applies both Ordinary Least Squares Regression and Overlap Weighting Analysis, which are reinforced through a Random Forest Regression Methodology for carrying out a robust cause-and-effect/correct prognosis.
5.1 Causal effect of digital financial literacy on financial behaviour
The econometrics show that the highest tertile of DFL has a significant positive impact on DFP relative to the lowest tertile. Though the baseline OLS estimate is fairly large at 0.857, the estimate is 1.305 for PSM and 1.412 using overlap weighting. The successive strengthening of estimates using different methods of identification is a good indicator of larger marginal effects of DFL on DFP after adjusting for both selection and observed characteristics.
These results are supported by the TPB, since it postulates that people are motivated by their own knowledge and perceived control to form an intended behaviour and perform specific actions. As stated by Ajzen (1991), higher levels of literacy shall lead to higher degrees of perceived competence, which will make digital participation even more possible. In conclusion to the HT, literacy improves skills and creates the ability to exploit digital tools in a more productive and safe manner (Becker, 1964).
The empirical findings emanating from contemporary research support this tenet, as, for example, Yadav et al. (2025a, b) have discovered positive outcomes emanating from DFL regarding the adoption of online payments and financial trust among Indian citizens, and, regarding the relationship between literacy and the use of digital wallets and mobile banks among all levels of society, Pattnayak and Sahoo (2024) offer proof. This study strengthens prior correlational evidence by showing results consistent with treatment effects after observable adjustments.
ML validation further extends this confidence in these results. The Random Forest model described approximately 68.7% of the variation in DFP with low error (Rˆ2 = 0.687, RMSE = 0.523), ensuring high predictive ability. Among these results, variable importance showed that “DFL” is the most important predictor (%IncMSE = 47.6%), followed by “education”, “age” and “income”; however, “gender” showed the least contribution. The partial dependence plot of “DFL vs. DFP” showed that it is monotonically positive. This means that as literacy increases, digital financial participation also increases. This consistency in econometric treatment-effect estimates and predictive validation from ML models further extends this validity to this inference in this study.
5.2 Gender-based heterogeneity and the empowerment effect
Another significant finding of the study is the gender-stratified result of the PSM, which captures the fact that the greater the DFL, the improved outcomes for both males and females, but the difference in impact is greater in the case of females as compared to males (ATT = 1.61 for women vs. 1.31 for males). This finding reveals the “leverage effect”, which states that while women have poorer starting levels of literacy, small gains in literacy result in bigger gains in behaviour.
This differential pattern implies that, based on the CAP Approach outlined by Sen (1999), the process of literacy could have a positive impact on the issue of agency, autonomy or the capability to convert resources to function. For women, the patriarchal system will make digital literacy important to counter such a system. As claimed by Mishra et al. (2024), recent evidence indicates that the DFL increases the financial confidence level of women and decreases their reliance on the intermediary system in the Indian context.
Cumulatively, these findings indicate that there is a grave issue of gender inequality in financial inclusiveness through digitisation. A report titled “Global Findex (2023)” released by the World Bank indicates that there is a 6–8 percentage point gap between men and women regarding formal account and mobile money holdings. The findings of this study indicate that such inequality can be brought down dramatically through literacy. In this way, DFL enables women to employ the use of technology as a financial tool of access to savings, credits and insurance, which is a traditional male domain.
5.3 Integration of econometric and machine-learning evidence
One of the methodological innovations of this study is to integrate causal inference and predictive modelling. Although econometric methods provide average treatment effects and identify causal directionality, ML algorithms can capture important non-linearities and interactions between predictors. In so doing, the study combines the interpretability of traditional econometrics with the predictive power of data-driven analytics, a framework now increasingly advocated in behavioural finance.
The complementary evidence from the two approaches suggests that DFL is not only causally related to digital financial engagement but also its strongest predictor, even when considering non-parametric relationships. A high explanatory power of the Random Forest model implies that the observable behavioural outcomes were substantially caused by measurable socio-demographic and literacy factors. This reduces unobserved heterogeneity. This multi-method validation adds external robustness and credibility to the conclusions of this study, something that single-method studies often lack.
5.4 Theoretical contributions
This research makes three major contributions to theory. First, the study empirically extends human capital theory into the digital domain, demonstrating that the investment in cognitive and technological literacy pays behavioural dividends. Second, it operationalises the theory of planned behaviour through the quantification of the way literacy influences behavioural intention and actions within the digital financial context. Third, it is consistent with capability theory, as it shows that literacy extends the set of financially attainable functions, especially for women and low-income groups.
Moreover, this study extends the TAM of Davis (1989) and its subsequent extensions, which suggest that perceived ease of use and usefulness are the basic drivers of technology adoption. DFL reinforces both perceptions: once people understand digital tools, they find them both easier and more useful, and hence use them more. These theoretical integrations collectively suggest that literacy operates as a multi-dimensional enabler, simultaneously cognitive, behavioural and sociostructurally.
6. Policy and practical implications
The results of this study have clear policy implications in terms of financial inclusion, designing fintech and capability-building initiatives. Although country-level policies like Digital India and PM Digital Abhiyan have made huge strides towards improving the infrastructure, it emerges clearly that literacy and not access, emerges as the limiting factor when it comes to adopting digital financial services. This means that there has to be a move from an infrastructure focus to a capability focus in designing financial inclusion policies wherein DFL is recognised as a major policy tool, rather than a peripheral one.
Since the effect of DFL on behavioural variables was found to be much stronger among women respondents, this aspect needs special attention while formulating intervention policies in favour of women. Policymakers may design programs around DFL through the integration of DFL courses with SHGs, the onboarding of Jan Dhan accounts, and community-level education initiatives.
The analysis is equally informative about the significance of socio-economic elements such as education, income and age in the development of DFL outcomes. As a result, segmented approaches to program implementation should be adopted. For example, the introduction of financial literacy in school education programs can facilitate early skill-building, while digital skills improvement for the elderly can compensate for the differences in the adoption process. Improving access through subsidising phones and cheaper Internet connections among low-income people may serve as an effective supplement to financial literacy initiatives.
In a data-centric approach to policymaking, the application of the Random Forest algorithm represents the possibility of identifying high-impact target audiences based on ML techniques. The model makes it possible to concentrate resources on individuals having low levels of literacy yet high responsiveness to behaviour change prompts.
Regarding financial institutions and fintech companies, the importance of designing literacy-friendly services is highlighted by the results. Designing simple interfaces and offering straightforward information may be accompanied by micro-learning features to increase consumer skills gradually. Transaction support, notifications of fraud and contextual assistance can help reduce risk perception among users.
At the policy level, there is a requirement to establish standardised models for assessing and tracking DFL in accordance with international standards such as OECD (2022). This will allow for improved monitoring of achievements and informed policymaking decisions.
6.1 Wider implications and future outlook
The combination of causal inference and ML in this article shows that the influence of DFL in achieving digital financial inclusion cannot be considered a complement but rather a basis. It shows that simply relying on technological advancement will not result in widespread use of technology in financial service provision without sufficient knowledge among the population. There is thus a need for changing the perspective of financial inclusion policy towards empowerment in the era of digitised finance.
With the gradual change of financial services in a phygital form, which combines physical and digital technologies, the importance of DFL will continue to grow because it is the determinant for who becomes beneficiaries of digitisation and who gets excluded. In addition, the identified gender differences show that DFL can be used not only as a means of economic growth but also for social justice through financial inclusion of women.
However, apart from the above findings, the study raises some interesting areas for future research. First, because of the cross-sectional nature of the data used in the research, it may not be possible to conclusively establish any temporal relationship between factors and behaviour; hence, future studies should consider longitudinal or panel datasets. Second, future studies should also adopt experimental or quasi-experimental design (for instance, randomised control trials). Third, the framework can be extended in cross-national settings, especially in developing countries, to test generalisability.
Fourth, incorporating behavioural theories in relation to aspects such as trust, risk perception and cybersecurity will improve our understanding of DFL-related behaviour. Lastly, future studies should investigate the possibility of using DFL in enhancing sustainability and responsibility in financial decision-making processes.
7. Limitations and future research directions
Although the contributions of this study are substantive, the following limitations should be acknowledged.
Causal identification in this study depends on observational data and matching methods that only account for observable confounders. Unobserved heterogeneity, like innate risk aversion or digital anxiety, may still bias results.
The cross-sectional design does not allow for an assessment of temporal dynamics; future longitudinal or panel data can capture whether literacy gains result in sustained or decaying digital behaviour over time.
While ML improves predictive robustness, it does not identify underlying causal pathways; integrating causal forests or uplift modelling could provide more granular insights into subgroup responsiveness.
The gender analysis, while revealing, can be deepened by qualitative exploration of mechanisms, like confidence, autonomy and peer effects, that could mediate women’s behavioural change.
Finally, while diverse, the data are limited to India; replication in other emerging economies could further strengthen generalisability and inform comparative policy learning.
8. Conclusion
This study confirms that DFL is indeed a crucial factor in influencing positive changes in DFP within India. The present study, adopting a multi-layered approach, combines covariate-adjusted regression, PSM, overlap weighting and machine-learning prediction, hence giving evidence that, although earlier studies found correlational evidence, DFL is strongly associated with improved digital financial behaviour and estimated treatment effects remain consistently positive across methods. Consistency of results across statistical and algorithmic models strengthens internal validity and places DFL as the most influential determinant of digital financial engagement, outweighing demographic attributes such as gender, age, education and income.
A key insight arising from heterogeneity analysis is that the marginal effect of DFL is higher for women, suggesting that literacy holds promise for bridging chronic gender gaps in digital finance adoption. What these points point to, in policy terms, is that without the cognitive and confidence-building element which literacy can provide, the behavioural uptake of digital finance will be constrained even as infrastructure, access and fintech innovation continue to advance. Therefore, improvement in DFL, especially for females, lower-income households and older segments of the population, will have disproportionately high returns in terms of inclusion, empowerment and usage depth.
The study also has implications for the methodological frontiers related to digital finance because it illustrates the value of combining causality and ML as tools for combining complementary forms of evidence. There is an excellent template to be replicated across related fields of study, such as financial inclusion, behavioural economics and technology adoption.
In general, however, the evidence reiterates the importance of the following strategic truth: “Digital finance is not inclusive by design-it is inclusive through capability. This means that the fastest payment infrastructure networks and the real-time settlement systems can be created in each of these countries. However, if their citizens’ knowledge and understanding are not built regarding their utilisation, digital finance remains poised to retain the same problems in their fold, irrespective of the claims of digitalisation to rectify them.” DFL is an ancillary element of Digital Development-it remains the behavioural backbone of DFL.
Future progress in digital economies will depend as much on how well systems are designed as on how well people are prepared to use them. DFL, in that sense, is more than just a policy variable; it acts as a bridge between access and agency, between technology and trust, and eventually between digitisation and real inclusion.
Ethical approval
Participants in this study, centred on the experiences and perspectives of individuals in the Delhi National Capital Region, affirm their comprehension of the research’s objectives.

