This paper aims to re-examine the uncovered interest parity (UIP) hypothesis, which posits efficiency in forward foreign exchange and rational expectations. Testing these assumptions involves estimating parameters in a k-step-ahead forecasting model, where forecast errors are expected to be serially correlated up to lags k 1. When errors are correlated beyond these lags, OLS is no longer consistent unless the regressors are exogenous.
The authors extend the FGLS procedure developed in Perron and González-Coya (2022) to a setting in which lagged dependent variables are included as regressors. The authors thus provide a consistent and efficient framework to estimate the parameters of a general k-step-ahead linear forecasting equation. Following the work of Perron and Olivari (2023), the authors introduce an instrumental variable (IV)-based approach for this problem that requires pre-determined but not necessarily exogenous IVs for consistency.
The authors apply the authors’ FGLS procedures to the analysis of the two main specifications to test the UIP. Contrary to most empirical results available in the literature, in particular those based on some OLS regression or GMM, the authors’ robust and efficient procedure cannot reject the null hypothesis that the UIP holds.
Overall, this study’s results can be viewed as overturning the so-called forward discount anomaly. The methods proposed can also be applied to a wide variety of contexts.
1. Introduction
In this paper, we re-examine the hypothesis of uncovered interest parity (UIP), which in its basic form implies that the (nominal) expected return to speculation in the forward foreign exchange market conditional on available information should be zero. This is an “efficient-markets hypothesis” (EMH) for foreign exchange markets: if all available information is used rationally by risk-neutral agents in determining the spot and forward exchange rates, then the expected rate of return to speculation will be zero and the foreign exchange market is said to be efficient. This is a joint hypothesis since it includes the assumption of rational expectations (REH) and the assumption that the risk premium for the forward rate is zero. In fact, rejection of the UIP hypothesis does not immediately translates into a rejection of the efficiency of the foreign exchange market, which could be due economic agents being risk averse. Still testing whether the UIP hypothesis holds has been and continue to be a topic of considerable interest from both theoretical and empirical perspectives.
Testing the rationality hypothesis and exchange market efficiency is embedded in the general problem of estimating the parameters of a k -step-ahead linear forecasting equation. When the sampling interval is finer than the interval over which forecasts are made (in this case the maturity time of the forward exchanges rates), the forecast error is serially correlated. As noted by Hansen and Hodrick (1980), under rational expectations (REH) the forecast error is serially correlated up to lag and OLS remains consistent but appropriate modifications in the estimation of the asymptotic covariance matrix are needed. However, if the REH is rejected and the forecast error is serially correlated beyond lag , OLS is no longer consistent when the regressors are not exogenous; see Perron and González-Coya (2022). Contrary to what is asserted in Hansen and Hodrick (1980), GLS is consistent when the regressors are pre-determined provided the roots of the MA polynomial are inside the unit circle, i.e. MA process is invertible, as shown in Perron and González-Coya (2022). Moreover, GLS remains consistent when the forecast error follows a linear invertible process.
The first contribution of this paper is to provide a consistent and efficient framework to estimate and perform tests of the parameters of a k -step-ahead linear forecasting equation that remains valid whether the REH holds or not. We apply the FGLS procedure developed in Perron and González-Coya (2022) which is consistent using non-exogenous regressors, provided the errors follows an stationary invertible linear process. The second contribution is to extend their FGLS procedure to cases with lagged dependent variables included as regressors. Following the work of Perron and Olivari (2023), we introduce an instrumental variable (IV)-based approach for this problem that requires pre-determined but not necessarily exogenous IVs for consistency. The third contribution is to apply our FGLS procedures to the two main regressions suggested in the literature to test the UIP. We use 30 years of data for three currencies, and we reconsider the framework and regressions used by Fama (1984) and Hansen and Hodrick (1980). We provide extensive simulation experiments to assess the finite sample performance of our FGLS procedure relative to OLS. We show that FGLS achieves important reductions in mean-squared error (MSE) and allow tests with much greater power.
The Fama regression assesses whether the current forward-spot differential, , is a good predictor of the future change in the spot rate, . Most results available in the literature suggest a negative estimate of the relevant parameter, which instead should take value one if the UIP holds. This is often referred to as the “forward discount anomaly”, which refers to the widespread empirical finding that the returns on nominal exchange rates is negatively correlated with the lagged forward premium. It implies an appreciating currency for the high interest rate country. It is an “anomaly” as rational expectations would imply the opposite; if all currencies are equally risky, investors would demand higher interest rates on currencies expected to fall in value. Our results, in contrast, indicate positive values, sometimes not significantly different from one. This finding suggests that the “forward discount anomaly” might be a consequence of OLS providing an inconsistent estimate and our FGLS procedure being consistent and efficient under a broader range of possible scenarios.
The regression adopted by Hansen and Hodrick (1980) is to test whether past values of the forward-spot differential help predict the current value , conditioning on some covariates involving the past forward-spot differentials from some other countries. Under the UIP and EMH, there should be no predictive power as all information contained in the information set at time t should already have been accounted for by the market in setting the forward rates. Here, contrary to most empirical results available in the literature, in particular those based on some OLS regression or GMM, our robust and efficient procedure cannot reject the null hypothesis that the UIP holds. Hence, overall, our results can be viewed as overturning the so-called “forward discount anomaly.”
The main methodological contribution of this paper is that it extends Perron and González-Coya (2022)FGLS procedure for the case in which lagged dependent variables are included as regressors (subsection 3.1). In subsection 4.2, we discuss an application of Perron and Olivari (2023) FGLS-IV method when the instrumental variable set includes lagged pre-determined regressors as instruments. These contributions, provide a consistent and efficient framework to estimate and perform tests of the parameters of a k-step-ahead linear forecasting equation that remains valid whether the REH holds or not. Empirically, the most significant contribution is that by applying the feasible GLS and GLS-IV methodologies described above to the estimation of the Fama and Hansen and Hodrick (1980) regressions, respectively, we obtain results that widely differ from the consensus in the literature. For the Fama regression, most results available in the literature suggest a negative estimate of the relevant parameter, often referred to as the “forward discount anomaly.” Our results in contrast, indicate positive values, sometimes not significantly different from one. This finding suggests that the “forward discount anomaly” might be a consequence of OLS providing an inconsistent estimate and our FGLS procedure being consistent and efficient under a broader range of possible scenarios. We also show statistically significant discrepancies between the OLS and GLS-IV estimates in the Hansen and Hodrick (1980) regression.
The remainder of this paper is as follows. Section 2 describes the formulations of the UIP hypothesis and reviews previous econometric tests. Section 3 describes the FGLS procedures and Section 4 introduces two instrumental variable (IV) based approaches for a setting in which lagged dependent variables are included as regressors. Section 5 studies the Fama (1984) regression and Section 6 is focused on the Hansen and Hodrick (1980) regression. For both specifications, we provide extensive Monte Carlo experiments to assess the finite sample performance of our FGLS procedures relative to OLS. We also provide empirical results for the sample period January 1993 to January 2023. In Section 7, we present a test for the hypothesis that the OLS residuals exhibits serial correlation of order greater than and discuss its empirical implications to understand the various conflicting results. In Section 8, we provide practical recommendations and brief concluding remarks are presented in Section 9. A supplement provide additional details.
2. Uncovered interest parity and the efficient market hypothesis
Let and , where and are the levels of the spot exchange rate and the k-period forward exchange rate at time t. With , the domestic nominal interest rate and the corresponding foreign interest rate, the theory of UIP implies that . Hence, UIP requires the twin assumptions of rational expectations and a constant or zero risk premium. Given the no arbitrage condition, the covered interest parity (CIP) condition implies that holds as an identity. Hence, the UIP condition is also frequently expressed as . Since is an approximate measure of the rate of return to speculation, we can express the efficient-markets hypothesis as . This implies forecast errors uncorrelated with information available at time t, .
2.1 Econometric tests
As in Hansen and Hodrick (1980), we consider the general problem of estimating the parameters of a k-step-ahead linear forecasting equation, . Then:
where rational expectations impose a specific structure on the forecast error . Due to the period overlap in the sequential k-step-ahead forecasts, has an MA() representation. Thus:
and the OLS estimates of are consistent since the regressors are pre-determined via the rational expectations hypothesis. We shall label this as the “RE case”. Note, however, that we could well be faced with a model of the form:
where is some serially correlated process involving innovations dated before period t; e.g. an process of the form for some sequence of i.i.d. innovations . In this case, if the regressors are non-exogenous with respect to past values of , OLS is no longer consistent since . We shall label this as the “general case”, as it encompasses the “RE case”. This motivates the necessity of an estimate that is consistent under the both the “RE case” and the “general case”, i.e. allowing the errors to be serially correlated beyond lag . Contrary to what is asserted in Hansen and Hodrick (1980), GLS is consistent in both cases, provided the errors can be represented as some invertible linear process, as shown in Perron and González-Coya (2022).
Example 1. A k-step ahead forecast error process can be serially correlated beyond lagsunder, e.g. Adaptive Expectations (AE) (seeMuth, 1960). An exponentially weighted moving average forecast arises from the following model of expectations adapting to changing conditions. Under AE, it is assumed that the forecast is changed from one period to the next by an amount proportional to the latest observed error:
As shown inMuth (1960), the solution of this difference equation is an exponentially weighted forecast . The forecast error is thus:
To characterize the process of the forecast error , we shall impose a functional form on . It is standard in the AE literature, to assume that has a permanent and a transitory component, where the permanent component is defined by with, and independent. In this case, the forecast error follows an process. The details of this derivation are spelled in the Appendix A.1.
Several tests of the UIP hypothesis have been proposed in the literature. Bilson (1981) and Fama (1984) analyzed regression (1) with and with . In this setting a test of UIP is that and . It has been noted by Fama (1984) and many subsequent studies that the estimated slope coefficient is frequently negative. This is known as the Forward Premium Anomaly: the country with the higher rate of interest has an appreciating currency rather than a depreciating currency; a violation of UIP. Hansen and Hodrick (1980) estimate the model (1) with , the forecast error, and uses lagged dependent variables as regressors; with . The test is and . A simplified version of this regression, with and test and , has been studied by Baillie et al. (2023). Hansen and Hodrick (1980) also proposed a test that includes the lagged forecast errors of four other currencies:
where is the forecast error for country i and is the forecast error for country . We focus on regression (2) as Hansen and Hodrick (1980) concludes that the multicountry test appears to be more powerful.
3. Feasible GLS
We consider the linear regression:
where the error term follows a short-memory linear process:
where . The polynomial is assumed to satisfy (a normalization), so that the process is short-memory and is invertible, i.e. we can write . We use the FGLS procedure developed in Perron and González-Coya (2022) to estimate regression (3). The idea is to consistently approximate using an autoregression of order :
with and as , to ensure consistent estimates; see Berk (1974). Replacing equation (3) and re-arranging terms we have the following:
with for . Equation (5) is often called the Durbin regression (see Durbin, 1970). The order of the autoregression, , is determined via the minimization of the BIC suggested by Schwarz (1978) for where is such that as . We use the method suggested by Ng and Perron (2005) to ensure a proper comparison across models with different values of , i.e. using the same effective number of observations. We estimate the Durbin equation (5) via OLS with . Using the OLS estimates of , , we construct the quasi-differenced variables:
The FGLS estimate of is the OLS estimate of the quasi-differenced regression:
The resulting estimate will be consistent provided the regressors are pre-determined with respect to values prior to period ; see Perron and González-Coya (2022).
Remark 1. The consistency of our FGLS and GLS-IV estimators is unaffected by conditional heteroskedasticity. The key requirement for consistency is that the error process admits a stationary invertible linear representation, so that the autoregressive filtering recovers serially uncorrelated errors and the quasi-differenced regressors remain pre-determined. Our variance estimators and the analogous IV-based expressions to be discussed below assume unconditional homoskedasticity of the quasi-differenced errors. Under conditional heteroskedasticity, these estimators remain consistent for the unconditional variance of the coefficient estimates.
3.1 Feasible GLS with lagged dependent variables
Consider regression (3) with for some , where is a vector of pre-determined regressors that includes a constant term. For simplicity we omit the constant term without loss of generality. Write the regression as:
If is autocorrelated beyond lags , OLS applied to regression (8) is not consistent as and are not independent. Wallis (1967) and Malinvaud (1966) studied regression (8) with where the error terms follows an process. Expression for the asymptotic bias of the OLS estimates of and are given by Malinvaud (1966) and Griliches (1961). We can write the Durbin regression as follows:
where , , for ; , for ; and for , . For the case and a scalar (i.e. ), Wallis (1967) and Malinvaud (1966) (p. 469) propose to estimate regression (9) using OLS and then estimate using . For the general case and , the same approach can be applied. Note that we need , otherwise the estimates are not consistent. But this condition will be satisfied, at least in large samples, when using the BIC to select the lag order. However, note that when , cannot be uniquely identified. We shall propose two estimation methods; one based on the Durbin regression (9) and one based on a first-stage instrumental variable (IV) estimate. These follow similar steps as in the feasible GLS procedure discussed above.
Remark 2. Estimating the autoregressive coefficients, for, from regression (9) using the Wallis (1967) and Malinvaud (1966) method does not allow us to uniquely identify in the general case with , and . However, we can obtain efficient estimates , by using a convex combination of the estimates () for . For ease of the exposition, suppose that . Then, we can construct efficient estimates , where the optimal that minimize is (the details are in the Appendix A.2):
In practice, we can estimate using a first order Taylor expansion:
The first order Taylor expansions of are cumbersome. Note that if the instruments are independent of each other, will be arbitrarily small in large samples. Hence, we set .
While the adaptive expectations framework provides an analytically tractable illustration of how forecast errors can be serially correlated beyond lag , we emphasize that it is one instance of a broader class of departures from the rational expectations paradigm that generate such excess autocorrelation. Several economically motivated models of expectations formation share this property. Under sticky information models Mankiw and Reis (2002), agents update their information sets infrequently, so that aggregate expectations adjust sluggishly to new information. This generates forecast errors that inherit the persistence of the underlying fundamentals, producing autocorrelation well beyond the lags implied by overlapping forecasts alone. Similarly, noisy information models in the tradition of Sims (2003) and Woodford (2003), in which agents observe fundamentals with noise and must solve a signal extraction problem, produce forecast errors whose serial correlation structure reflects the dynamics of the signal-to-noise ratio rather than the forecast horizon alone. Models with heterogeneous expectations – where agents use different forecasting rules and the population weights evolve over time – also generate aggregate forecast errors with rich autocorrelation structures; see, e.g. De Grauwe and Grimaldi (2006). Finally, behavioral models featuring extrapolative or momentum-based expectations Barberis et al. (1998) can produce persistent forecast errors when agents systematically over- or under-react to recent exchange rate movements. The key insight is that our FGLS procedure does not require the econometrician to specify or identify which of these mechanisms is operative. It requires only that the error process be representable as a stationary invertible linear process. This agnosticism is a practical advantage: the procedure delivers consistent and efficient estimates regardless of whether the excess serial correlation originates from adaptive learning, rational inattention, heterogeneous beliefs or any other mechanism that generates a departure from the pure MA structure.
4. Instrumental variables
An alternative to estimate regression (8) with is to use an instrumental variable procedure. If the regressors are exogenous, then are valid instruments for and the two-stage least squares (2SLS) estimates will be consistent. Liviatan (1963) propose to use as instruments for and as an instrument for itself, so that the instrument set is . We can potentially select a larger set of instrumental variables for ; e.g. lags or order of . In this case , for some . Note that p can be adaptively selected using appropriate tests for over-identifying restrictions (see Small, 2007). However, if the regressors are not exogenous and is autocorrelated beyond lags , are no longer valid instruments. Hence, the IV estimate of regression (8) using the instrument set will not be consistent. This motivates the use of the GLS-IV procedure suggested by Perron and Olivari (2023). The idea is simple: first transform the model to have serially uncorrelated errors and then estimate the transformed model via IV using the set of transformed instruments. The resulting estimate will be consistent as the instrument set and regressors are pre-determined in the transformed model.
4.1 Estimator valid with exogenous instruments
We first discuss two methods that are valid with exogenous instruments if is autocorrelated beyond lags . One is the widely used so-called “optimal GMM” procedure. The other is akin to the GLS procedure discussed above but with the first-step using the GMM estimate to obtain estimate of the residuals and construct the autoregressive filtering.
4.1.1 GMM.
Using the set of exogenous instruments , then, as shown in Hansen (1982), the best estimator of based on the instruments and the moment condition is as follows:
where the matrix is given by . We can write , where and . Then can thus be consistently estimated by , where , with , is some kernel or weight function and m is the bandwidth; see, e.g. Andrews (1991).
4.1.2 GMM-GLS-IV.
We also consider a FGLS method that does not rely on the Durbin regression (9). Instead, it uses a first-stage GMM estimate to obtain a consistent estimate of the residuals, . We can thus identify the autoregressive parameters when the instruments are exogenous. The GMM-GLS-IV procedure to estimate regression (8) with exogenous instruments is the following: 1) obtain the GMM estimator of regression (8), given by (10) using the set of instruments . Compute the residuals, ; 2) select the order of the autoregression:
via the minimization of the BIC for where is such that as ; 3) estimate the autoregression (11) with to obtain consistent estimates (); 4) use () to construct the quasi-differenced variables , and ; 5) the GMM-GLS-IV estimate of is the IV estimate of the quasi-differenced regression:
using the set of quasi-differenced instruments .
4.2 Non-exogenous instruments: GLS-IV
We next describe the GLS-IV procedure to estimate regression (8), which is valid with non-exogenous instruments, provided they are pre-determined. The steps are the following: 1) Select the order of the Durbin regression (9), via the minimization of BIC for where is such that as ; 2) Estimate the Durbin regression (9) with the selected value using OLS. The estimates the autoregressive coefficients, , are obtained using the efficient method described in Remark 2; 3) Use , to construct the quasi-differenced variables , and , as in Step 4 for GMM-GLS-IV; 4) The GLS-IV estimate of is the IV estimate of the quasi-differenced regression (12) using the set of quasi-differenced instruments .
Given that the OLS estimates from the Durbin regression (9) are consistent and that the transformed model (12) has serially uncorrelated errors, the GLS-IV estimates are consistent under the stated conditions. To the best of our knowledge, there is no other consistent estimation method requiring only pre-determined regressors for the general linear regression (8) without restricting the error process . Maximum likelihood and non-linear least squares require an a priori known error process. In the simulation experiments reported in subsection 6.1.2, we show that the fact that is not uniquely identified when we have more than one regressor (), does not affect the efficiency of GLS-IV.
5. Fama regression and the forward discount anomaly
The Fama (1984) regression is as follows:
with for 1-month forward rates and for 3-month forward rates, when using weekly data. Estimates of (13) tell us whether the current forward-spot differential, , has power to predict the future change in the spot rate, . Evidence that is significantly different from zero means that the forward rate observed at t has information about the spot rate to be observed at . Under the efficient-market hypothesis, we have . Under the null hypothesis, the log of the forward rate provides an unbiased forecast of the log of the future spot exchange rate. Derivations from are sometimes interpreted as a measure of the variation of the premium in the forward rate.
Let be the OLS estimate of in regression (13). If the estimator is consistent, we have as follows:
If expectations are rational, then , where is the forecast error. In this case, . The foreign exchange risk premium when expectations are rational is defined as . Under risk neutrality, expected profits from forward market speculation would be zero as agents would drive into equality with . Write and replace into equation (14), so that , where:
A negative estimate of in regression (13) is a robust finding in the literature (see Engel, 1996). This is known as the “forward discount anomaly”; it is a widespread empirical finding that the returns on nominal exchange rates appear to be negatively correlated with the lagged forward premium. Bilson (1981) and Fama (1984) provide evidence that the estimates of are less than zero. Many subsequent studies have confirmed that finding, for dollar exchange rates and a large number of exchange rates and time periods (see, for example, Bekaert and Hodrick, 1993; Backus et al., 1993; Hai et al., 1997). Froot (1990) notes that the average value of over 75 published estimates is −0.88. Only a few of the estimates are greater than zero, and none is greater than 1. The forward discount anomaly implies an appreciating currency for the high interest rate country. It is an “anomaly” as rational expectations would imply the opposite; if all currencies are equally risky, investors would demand higher interest rates on currencies expected to fall in value. The survey by Engel (1996) focuses on the possibility that among the possible explanations for finding . Other possible interpretations are that the forward rate is a biased predictor of the future spot rate, and/or that it is evidence of a time-varying risk premium. In this paper, we provide a novel insight on this empirical regularity. We argue that the OLS and GMM estimates of are not consistent, as the error term in regression (13) is serially correlated beyond lags . We find that the FGLS estimates, which are consistent under a broader range of conditions, are significantly non negative but in general smaller than 1.
5.1 Data set
Daily data were obtained for the spot exchange rates for the UK pound (US-UK), Canadian dollar (US-CAD) and Japanese yen (US-JP) as well as the one- and three-month forward exchange rates data for the three currencies. As in Hansen and Hodrick (1980), the data were sampled to form a weekly series constructed by taking observation on Tuesday of each week. If no Tuesday observation was available, we used the Wednesday observations. The source of the forward exchange data for US-UK and US-CAD is Barclays Bank PLC; the source for US-JP is the Bank of Tokyo Mitsubishi. For all the data sets, we use all the information available until January 17, 2023 but the starting date of the time series differ: a) US-UK: starting date of October 11, 1983, 2,050 observations; b) US-CAD: starting date of December 14, 1984, 1,989 observations; c) US-JP: starting date of September 1, 1993, 1,535 observations.
5.2 Simulation design
In this section, we provide simulation results related to the Fama (1984) regression (13) under the “RE case” or efficient market hypothesis (EMH), . We consider for one-month forward rates and for three-month forward rate. As discussed in Section 2, the EMH implies that has an representation. We simulate an process with parameters calibrated to replicate the observed autocorrelation function up to lag of the FGLS residuals of equation (13) using US-UK data. The details are presented in the Appendix A.3. For we have an representation with:
For we have an representation with coefficients:
We simulate an error process based on equation (15), i.e. , where and is estimated using the residuals of an initial Fama FGLS regression (13) for US-UK. The data generating process (DGP) uses observed in the data for US-UK, and the stated simulated error process. We artificially generate to satisfy the null hypothesis . Thus, the DGP is as follows:
We also consider a departure from the “RE case” with errors following the “general case” with serial correlation at lags . We assume that follows an process with MA coefficients given by (15) and (16) depending on and , respectively. The AR coefficient is set to . We consider two sampling periods; the complete sample from October 11, 1983 to January 17, 2023 with 2,050 observations and the one spanning November 1. 1989 to April 1, 2021, as considered in Baillie et al. (2023). We perform 5,000 replications. We consider the following estimators of , from regression (13): a) The OLS estimate with HAC standard errors based on the quadratic spectral window suggested by Andrews (1991) with automatic bandwidth selection using an approximation; b) The Durbin estimate based on the regression (5) with and and selected using the BIC. For the complete sample, we consider and for the sub-sample we set ; c) The FGLS estimate based on regression (7) with , and selected using the BIC. The same values of are used. We estimate the sample variance of the FGLS estimator using , where is the matrix of quasi-differenced regressors (including a constant term) and is the sample variance of the FGLS residuals, . The confidence intervals at the 95% nominal level for the jth coefficient are obtained using , where is the 0.975 quantile of the normal distribution.
The simulation results are presented in Table 1 for and Table 2 for . In line with the theory, the mean squared error (MSE) of OLS is small when the error term follows an process and deteriorates when the error is serially correlated at lags . Clearly, FGLS outperforms OLS and Durbin in all cases. Even under the “RE case”, the MSE of OLS is on average 3.4 times larger than of FGLS, while the MSE of Durbin is close to that of OLS case. When the errors follow an process, OLS is no longer efficient, and its MSE is on average 27 times that of FGLS. As expected, the Durbin estimate remains consistent in this case, but its MSE is on average three times that of FGLS. FGLS also has the smallest variance. The variance of Durbin is on average three times the variance of FGLS for and two times the variance of FGLS for . The coverage rates of the confidence interval for FGLS are near the nominal 90% and have shortest lengths. OLS with HAC standard errors exhibits substantial size distortions with errors, especially when case. The OLS based HAC standard errors provides confidence intervals close to the nominal level in some cases, at the expense of a very large variance. For the case with the variance of OLS is 30 (25) times the variance of FGLS. This results are in line with those from Perron and González-Coya (2022).
Simulation results under EMH, one-month forward rate. RE implies MA(3) errors; general implies ARMA(1,3) errors
| Period | Case | Estimator | MSE | Bias | Variance | Coverage | Length |
|---|---|---|---|---|---|---|---|
| 11/83–01/23 | RE | OLS | 0.8152 | 0.7304 | 0.7598 | 0.94 | 3.3952 |
| FGLS | 0.1801 | 0.3388 | 0.1919 | 0.96 | 1.7167 | ||
| Durbin | 0.4950 | 0.5553 | 0.5057 | 0.95 | 2.7888 | ||
| General | OLS | 4.9239 | 1.7974 | 4.4738 | 0.93 | 8.2112 | |
| FGLS | 0.1445 | 0.3068 | 0.1525 | 0.96 | 1.5311 | ||
| Durbin | 0.4927 | 0.0035 | 0.4955 | 0.95 | 2.7607 | ||
| 01/89–04/21 | RE | OLS | 1.0063 | 0.7929 | 1.0019 | 0.94 | 3.8879 |
| FGLS | 0.3108 | 0.4422 | 0.3347 | 0.95 | 2.2667 | ||
| Durbin | 0.9838 | 0.7883 | 1.0471 | 0.96 | 4.0135 | ||
| General | OLS | 6.2271 | 1.9730 | 5.8364 | 0.92 | 9.3351 | |
| FGLS | 0.2736 | 0.4212 | 0.2926 | 0.96 | 2.1200 | ||
| Durbin | 0.9924 | 0.7929 | 1.0502 | 0.95 | 4.0196 |
| Period | Case | Estimator | Bias | Variance | Coverage | Length | |
|---|---|---|---|---|---|---|---|
| 11/83–01/23 | 0.8152 | 0.7304 | 0.7598 | 0.94 | 3.3952 | ||
| 0.1801 | 0.3388 | 0.1919 | 0.96 | 1.7167 | |||
| Durbin | 0.4950 | 0.5553 | 0.5057 | 0.95 | 2.7888 | ||
| General | 4.9239 | 1.7974 | 4.4738 | 0.93 | 8.2112 | ||
| 0.1445 | 0.3068 | 0.1525 | 0.96 | 1.5311 | |||
| Durbin | 0.4927 | 0.0035 | 0.4955 | 0.95 | 2.7607 | ||
| 01/89–04/21 | 1.0063 | 0.7929 | 1.0019 | 0.94 | 3.8879 | ||
| 0.3108 | 0.4422 | 0.3347 | 0.95 | 2.2667 | |||
| Durbin | 0.9838 | 0.7883 | 1.0471 | 0.96 | 4.0135 | ||
| General | 6.2271 | 1.9730 | 5.8364 | 0.92 | 9.3351 | ||
| 0.2736 | 0.4212 | 0.2926 | 0.96 | 2.1200 | |||
| Durbin | 0.9924 | 0.7929 | 1.0502 | 0.95 | 4.0196 |
We use US-UK observed data , , . For OLS we use HAC standard errors as described in the text
Simulation results under EMH, three-month forward rate. RE implies MA(11) errors; general implies ARMA(1,11) errors
| Period | Case | Estimator | MSE | Bias | Variance | Coverage | Length |
|---|---|---|---|---|---|---|---|
| 11/83–01/23 | RE | OLS | 1.6898 | 1.0541 | 1.5036 | 0.92 | 4.7404 |
| FGLS | 0.7098 | 0.6761 | 0.7261 | 0.95 | 3.3387 | ||
| Durbin | 1.3941 | 0.9358 | 1.4148 | 0.95 | 4.6647 | ||
| General | OLS | 18.5744 | 3.4972 | 15.3260 | 0.91 | 15.0322 | |
| FGLS | 0.6019 | 0.6283 | 0.6332 | 0.96 | 3.1202 | ||
| Durbin | 1.3964 | 0.9370 | 1.4193 | 0.95 | 4.6720 | ||
| 01/89–04/21 | RE | OLS | 1.8895 | 1.0869 | 1.6558 | 0.91 | 4.9454 |
| FGLS | 0.9995 | 0.7817 | 0.9834 | 0.94 | 3.8838 | ||
| Durbin | 1.9843 | 1.1502 | 2.0973 | 0.96 | 5.6801 | ||
| General | OLS | 20.6180 | 3.5933 | 16.2122 | 0.88 | 15.3295 | |
| FGLS | 0.8825 | 0.7451 | 0.9027 | 0.95 | 3.7256 | ||
| Durbin | 1.9782 | 1.1467 | 2.0995 | 0.96 | 5.6832 |
| Period | Case | Estimator | Bias | Variance | Coverage | Length | |
|---|---|---|---|---|---|---|---|
| 11/83–01/23 | 1.6898 | 1.0541 | 1.5036 | 0.92 | 4.7404 | ||
| 0.7098 | 0.6761 | 0.7261 | 0.95 | 3.3387 | |||
| Durbin | 1.3941 | 0.9358 | 1.4148 | 0.95 | 4.6647 | ||
| General | 18.5744 | 3.4972 | 15.3260 | 0.91 | 15.0322 | ||
| 0.6019 | 0.6283 | 0.6332 | 0.96 | 3.1202 | |||
| Durbin | 1.3964 | 0.9370 | 1.4193 | 0.95 | 4.6720 | ||
| 01/89–04/21 | 1.8895 | 1.0869 | 1.6558 | 0.91 | 4.9454 | ||
| 0.9995 | 0.7817 | 0.9834 | 0.94 | 3.8838 | |||
| Durbin | 1.9843 | 1.1502 | 2.0973 | 0.96 | 5.6801 | ||
| General | 20.6180 | 3.5933 | 16.2122 | 0.88 | 15.3295 | ||
| 0.8825 | 0.7451 | 0.9027 | 0.95 | 3.7256 | |||
| Durbin | 1.9782 | 1.1467 | 2.0995 | 0.96 | 5.6832 |
We use US-UK observed data , , . For OLS we use HAC standard errors as described in the text
For the simulations related to power, we use US-UK observed data for the period November 1, 1989 to April 1, 2021 with one-month forward rates, . The DGP is as follows:
where . The null value is . The error process is simulated as before; under the “RE case” follows an process, and under the “general case” follows an process. Figures 1 and 2 present plots of the empirical rejection frequencies for t-test of with nominal size 0.05 for the OLS estimates with HAC- standard errors, the FGLS estimates based on regression (7) and the Durbin estimates based on the regression (5). For the latter two, is selected via BIC with . Figure 1 pertains to “RE case” with the errors an process, while Figure 2 pertains to the “general case” with errors following an process.
The line chart presents rejection frequency on the y-axis against beta on the x-axis for 3 statistical methods, F G L S, O L S, and Durbin. The F G L S curve begins at the highest rejection frequency near 0.58 at beta 0, decreases sharply to approximately 0.05 around beta 1, and then rises again to about 0.41 at beta 2. The O L S and Durbin curves show lower rejection frequencies throughout, both declining towards beta 1 before increasing gradually at higher beta values. The chart demonstrates a U-shaped trend for all methods, with F G L S exhibiting the greatest variation and highest rejection frequency overall.Empirical Rejection Frequencies Of Nominal 5% t-Test of . Fama regression, RE case with MA(3) errors
The line chart presents rejection frequency on the y-axis against beta on the x-axis for 3 statistical methods, F G L S, O L S, and Durbin. The F G L S curve begins at the highest rejection frequency near 0.58 at beta 0, decreases sharply to approximately 0.05 around beta 1, and then rises again to about 0.41 at beta 2. The O L S and Durbin curves show lower rejection frequencies throughout, both declining towards beta 1 before increasing gradually at higher beta values. The chart demonstrates a U-shaped trend for all methods, with F G L S exhibiting the greatest variation and highest rejection frequency overall.Empirical Rejection Frequencies Of Nominal 5% t-Test of . Fama regression, RE case with MA(3) errors
The line chart shows rejection frequency on the y-axis plotted against beta on the x-axis for F G L S, O L S, and Durbin statistical methods. The F G L S curve starts above 0.60 at beta 0, decreases steeply to its minimum near beta 1, and then rises again towards beta 2. The O L S curve remains comparatively stable between 0.07 and 0.11 across most beta values, while the Durbin curve declines from approximately 0.26 to around 0.05 before increasing moderately at higher beta values. The results indicate that F G L S is more sensitive to changes in beta than the other methods.Empirical rejection frequencies of nominal 5% t-Test of . Fama regression, General case with ARMA(1,3) errors
The line chart shows rejection frequency on the y-axis plotted against beta on the x-axis for F G L S, O L S, and Durbin statistical methods. The F G L S curve starts above 0.60 at beta 0, decreases steeply to its minimum near beta 1, and then rises again towards beta 2. The O L S curve remains comparatively stable between 0.07 and 0.11 across most beta values, while the Durbin curve declines from approximately 0.26 to around 0.05 before increasing moderately at higher beta values. The results indicate that F G L S is more sensitive to changes in beta than the other methods.Empirical rejection frequencies of nominal 5% t-Test of . Fama regression, General case with ARMA(1,3) errors
Note that in Figure 1, “RE case” with errors, OLS and Durbin has very small rejection frequencies even when is far from the null value 1, even though they are consistent, which can be attributed to their lack of efficiency in finite samples. As shown in Figure 2, the power functions are similar in the case with errors, with the exception that the power function of OLS decreases and flattens with an almost constant 10% rejection frequency for all values, despite having a more liberal size.
5.3 Empirical results
We present the estimation results of regression (13) using the three estimates considered before: OLS, Durbin estimates based on the regression (5) and FGLS based on regression (7). We use data for three currencies, US-UK, US-CAD and US-JP and we consider two sampling periods; the first one is the complete sample and the second one spans from November 1, 1989 to April 1, 2021, the sampling period considered in Baillie et al. (2023). For the Durbin and FGLS estimates we set for the full sample and for the sub-sample period. Tables 3 and 4 present the estimation results using one-month forward rates () and the three-month forward rates (), respectively. In line with the empirical regularity in the literature, the OLS estimates of are not significantly positive in all cases. In contrast, the Durbin and FGLS estimates are positive in most cases and significantly non negative in some. In some cases, such as the US-JP exchange rate for three-month forward rates, we observe a negative significant OLS estimate and a positive significant FGLS estimate. Recall that the Durbin and FGLS estimates are consistent even when the error follows a general linear process. Section 7 present the outcome of the Cumby and Huizinga (1992) test (CH-test) for autocorrelation at lags . The outcome is a rejection in all cases. Hence, our finding suggests that the forward discount anomaly might be a consequence of OLS providing an inconsistent estimate of caused by non-exogenous regressors. We thus interpret the large differences between the OLS and the FGLS estimates as evidence of OLS being inconsistent.
Fama (1984) model estimation results, one-month forward rate (SE in parentheses)
| FX | Period | OLS | Durbin | FGLS | OLS | Durbin | FGLS |
|---|---|---|---|---|---|---|---|
| US-UK | 11/1983–01/2023 | −0.00167 (0.001) | −0.00031 (4.1E-04) | −0.00006 (3.3E-04) | −1.08291 (0.316) | −0.09390 (0.380) | 0.08457 (0.191) |
| 10/1989–04/2021 | −0.00241 (0.002) | −0.00067 (0.001) | −0.00027 (0.001) | −1.32248 (0.359) | 0.44659 (0.781) | −0.28232 (0.464) | |
| US-CA | 12/1984–01/2023 | −0.00015 (0.001) | −0.00000 (2.5E-04) | 0.00004 (2.1E-04) | −0.24314 (0.333) | 0.14125 (0.376) | 0.20060 (0.212) |
| 10/1989–04/2021 | −0.00005 (0.002) | 0.00000 (0.001) | 0.00005 (0.001) | 0.04296 (0.384) | 0.49843 (0.652) | 0.19765 (0.375) | |
| US-JP | 09/1993–01/2023 | −0.00165 (0.002) | −0.00018 (0.001) | 0.00005 (3.3E-04) | −0.08744 (0.110) | 0.26966 (0.145) | 0.17845 (0.081) |
| 09/1993–04/2021 | −0.00101 (0.003) | −0.00027 (0.001) | 0.00011 (0.001) | −0.13273 (0.110) | 0.14601 (0.282) | 0.31225 (0.152) | |
| | | ||||||
|---|---|---|---|---|---|---|---|
| FX | Period | Durbin | Durbin | ||||
| US-UK | 11/1983–01/2023 | −0.00167 (0.001) | −0.00031 (4.1E-04) | −0.00006 (3.3E-04) | −1.08291 (0.316) | −0.09390 (0.380) | 0.08457 (0.191) |
| 10/1989–04/2021 | −0.00241 (0.002) | −0.00067 (0.001) | −0.00027 (0.001) | −1.32248 (0.359) | 0.44659 (0.781) | −0.28232 (0.464) | |
| US-CA | 12/1984–01/2023 | −0.00015 (0.001) | −0.00000 (2.5E-04) | 0.00004 (2.1E-04) | −0.24314 (0.333) | 0.14125 (0.376) | 0.20060 (0.212) |
| 10/1989–04/2021 | −0.00005 (0.002) | 0.00000 (0.001) | 0.00005 (0.001) | 0.04296 (0.384) | 0.49843 (0.652) | 0.19765 (0.375) | |
| US-JP | 09/1993–01/2023 | −0.00165 (0.002) | −0.00018 (0.001) | 0.00005 (3.3E-04) | −0.08744 (0.110) | 0.26966 (0.145) | 0.17845 (0.081) |
| 09/1993–04/2021 | −0.00101 (0.003) | −0.00027 (0.001) | 0.00011 (0.001) | −0.13273 (0.110) | 0.14601 (0.282) | 0.31225 (0.152) | |
Fama (1984) model estimation results, three-month forward rate (SE in parentheses)
| FX | Period | OLS | Durbin | FGLS | OLS | Durbin | FGLS |
|---|---|---|---|---|---|---|---|
| US-UK | 11/1983–01/2023 | −0.00439 (0.004) | −0.00035 (4.1E-04) | 0.00001 (3.3E-04) | −0.98614 (0.218) | 0.64187 (0.382) | 0.29696 (0.278) |
| 10/1989–04/2021 | −0.00011 (0.004) | −0.00011 (0.001) | 0.00018 (0.001) | 0.34508 (0.238) | 1.25038 (0.474) | 1.15847 (0.332) | |
| US-CA | 12/1984–01/2023 | −0.00036 (0.003) | 0.00004 (2.5E-04) | 0.00006 (2.1E-04) | −0.21093 (0.218) | 0.51211 (0.343) | 0.31376 (0.261) |
| 10/1989–04/2021 | −0.00040 (0.003) | 0.00002 (0.001) | −0.00002 (0.001) | 0.19149 (0.252) | 0.55918 (0.399) | 0.31524 (0.301) | |
| US-JP | 09/1993–01/2023 | −0.00355 (0.005) | −0.00017 (0.001) | −0.00015 (3.3E-04) | −0.10423 (0.138) | 0.45446 (0.152) | 0.30124 (0.099) |
| 09/1993–04/2021 | −0.00158 (0.005) | −0.00003 (0.001) | −0.00002 (0.001) | −0.26883 (0.139) | 0.49440 (0.157) | 0.32649 (0.101) | |
| | | ||||||
|---|---|---|---|---|---|---|---|
| FX | Period | Durbin | Durbin | ||||
| US-UK | 11/1983–01/2023 | −0.00439 (0.004) | −0.00035 (4.1E-04) | 0.00001 (3.3E-04) | −0.98614 (0.218) | 0.64187 (0.382) | 0.29696 (0.278) |
| 10/1989–04/2021 | −0.00011 (0.004) | −0.00011 (0.001) | 0.00018 (0.001) | 0.34508 (0.238) | 1.25038 (0.474) | 1.15847 (0.332) | |
| US-CA | 12/1984–01/2023 | −0.00036 (0.003) | 0.00004 (2.5E-04) | 0.00006 (2.1E-04) | −0.21093 (0.218) | 0.51211 (0.343) | 0.31376 (0.261) |
| 10/1989–04/2021 | −0.00040 (0.003) | 0.00002 (0.001) | −0.00002 (0.001) | 0.19149 (0.252) | 0.55918 (0.399) | 0.31524 (0.301) | |
| US-JP | 09/1993–01/2023 | −0.00355 (0.005) | −0.00017 (0.001) | −0.00015 (3.3E-04) | −0.10423 (0.138) | 0.45446 (0.152) | 0.30124 (0.099) |
| 09/1993–04/2021 | −0.00158 (0.005) | −0.00003 (0.001) | −0.00002 (0.001) | −0.26883 (0.139) | 0.49440 (0.157) | 0.32649 (0.101) | |
The economic intuition behind our findings can be summarized as follows. The forward discount anomaly – the robust negative estimate of in the Fama regression – has traditionally been interpreted as reflecting either a time-varying risk premium that covaries negatively with the forward discount, a failure of rational expectations or some combination of the two. Our results point to a fundamentally different explanation: the anomaly is primarily a statistical artifact arising from the inconsistency of OLS when the forecast error is serially correlated beyond lag . To see why, note that in the Fama regression the regressor is not strictly exogenous with respect to the full error process . When exhibits serial correlation at lags – a feature we confirm empirically via the Cumby–Huizinga test (Tables 10 and 11) – past realizations of u are correlated with current values of through their joint dependence on fundamentals that drive both the forward discount and subsequent spot rate movements. This induces a negative bias in the OLS estimate of , mechanically pushing it below zero. The FGLS procedure eliminates this bias by filtering the regression to produce serially uncorrelated errors, thereby restoring the orthogonality between the transformed regressors and transformed errors that is required for consistency. The resulting positive estimates of are thus not an anomalous new finding; they are what one should expect once the estimation procedure is robust to the actual structure of the error process. In other words, the forward rate does contain information about future spot rates broadly consistent with UIP, but this relationship is obscured in standard OLS regressions by an endogeneity bias induced by excess serial correlation in the forecast errors. The decades-long literature documenting negative estimates may thus have been attributing to economic forces – large and volatile risk premia, systematic expectational failures – what is more parsimoniously explained as a consequence of inconsistent estimation.
Our FGLS estimates of are positive across all currencies and forward rate maturities, but in most specifications they remain below unity. This pattern warrants discussion in terms of the risk premium. Recall that under rational expectations, , where captures the covariance between the risk premium and the forward discount , scaled by the variance of the forward discount. A consistent estimate of that is positive but less than one implies , which is consistent with a time-varying risk premium that covaries positively with the forward discount, which is the sign predicted by standard asset pricing theory. Under the consumption-based CAPM or its Epstein–Zin generalization, currencies with higher interest rates are expected to depreciate, but investors require a positive premium for bearing the associated consumption risk, so the forward rate exceeds the expected future spot rate by a modest margin. Our estimates suggest that this risk premium component, while present, is economically small. To illustrate, consider the US-JP estimates for three-month forward rates (Table 4): the FGLS estimate of implies , whereas the OLS estimate of implies . The latter would require the variance of the risk premium to exceed the variance of the expected depreciation – a condition that Fama (1984) himself noted was difficult to reconcile with plausible models of risk compensation. Our FGLS estimates, by contrast, imply a more moderate risk premium whose magnitude is broadly consistent with calibrated models of international asset pricing; see, e.g. Lustig and Verdelhan (2007) and Bansal and Shaliastovich (2013). We note, however, that our framework does not allow us to separately identify the risk premium and expectational components. The departure of from unity could reflect a genuine, economically meaningful risk premium, a residual (but small) departure from rational expectations or both. What our results do establish is that the magnitude of this departure is far smaller than previously documented, and in particular that the implausibly large negative estimates of from OLS are attributable to inconsistent estimation rather than to economic fundamentals.
6. Hansen and Hodrick regression
We now turn to an estimation problem that shares some of the main features, though with added complexities. Our aim is to efficiently estimate Hansen and Hodrick (1980) regression:
where and for . Note that if is significantly different from zero, then is correlated with the regression variable and can be potentially used as an instrument. If the lagged forecast error of country , , is exogenous, i.e. uncorrelated with the residuals for all t and s then GMM and GMM-GLS-IV are consistent. If is only pre-determined, the OLS and GMM (and thereby GMM-GLS-IV) estimates are not consistent in general, but GLS-IV will remain consistent. For GMM and the first-step for GMM-GLS-IV we consider the set of instrumental variables . We use weekly data for 1-month forward rates and 3-month forward rates. In the former case, whereas in the second . We consider the same data set as in subsection 5.1. We first present simulations tailored to this problem to shed light on the properties of the various estimators under a range of plausible scenarios.
6.1 Simulation results
We present two sets of Monte Carlo experiments. The DGP is inspired by the Hansen and Hodrick (1980) regression. We consider one-month forward rates so that :
where , and , . In the first set of Monte Carlo experiments, the simulation design is based on observed weekly spot and forward exchange rates for US-UK, US-CAD and US-JP. We use actual forecast errors for US-CAD and US-JP. By construction, the regressors will be exogenous. In subsection 6.1.2, we present the second set of Monte Carlo experiments, where the regressors are jointly simulated with the error process , so they are serially correlated and non-exogenous.
We start with the case of exogenous instruments. We consider the DGP (20) under the null hypothesis . We set , so that , can be used as instruments. We use actual observed data for , and we generate according with the DGP (20) imposing the null hypothesis. As discussed in Section 2, the “RE case” implies that has an representation. We simulate an process with parameters calibrated to replicate the observed autocorrelation function up to lag 3 of the GLS-IV residuals from equation (19) using the real data set, following the same procedure as described in Appendix A.3. The resulting parameters are as follows:
We simulate an error process based on equation (21), , where and is estimated using the GLS-IV residuals from the initial regression (20). We use the weekly spot exchange rates for US-UK, US-CAD, US-JP for the period between period November 2010 to April 2020 (492 observations). We perform 5,000 replications. We consider the following estimates: a) OLS applied to regression (20). We use HAC standard errors with the Quadratic Spectral weighting scheme of Andrews (1991) with automatic bandwidth selection using an approximation; b) GMM: the estimate (10) from regression (20) using the optimal weighting matrix as described in subsection 4.1.1. We use the set of instruments ; c) GLS-IV: the estimate from the procedure described in subsection 4.2. For the Durbin regression (9), we set . The autoregressive coefficients are estimated using the efficient method described in Remark 2. For the IV estimates, we use the set of quasi-differenced instruments . The variance estimate is as follows:
where is the matrix of quasi-differenced regressors (including a constant term) and is the sample variance of the GLS-IV residuals ; d) GMM-GLS-IV: the estimate from the procedure described in subsection 4.1.2. Step 1 uses the GMM estimate described above. For the autoregression (11), we set . For the IV estimate, we use the set of quasi-differenced instruments . The variance estimate is computed as:
where is the matrix of quasi-differenced regressors (including a constant term) and is the variance of the GMM-GLS-IV residuals:
6.1.1 The case with exogenous instruments.
We start with the case with exogenous instruments and first assess the finite sample size of the estimators. The simulation results are presented in the first panel of Table 5, which report the MSE, bias, variance, coverage rate and average length of confidence intervals for the parameter . Under the “RE case”, the error term follows an process so that the OLS estimates are consistent. The GMM estimate has a variance that is half of the OLS variance but with a coverage rate below the nominal level. The GMM-GLS-IV procedure using the GMM as a first step estimate achieves an important reduction in MSE while maintaining coverage rates near the nominal level. The finite sample performance of GLS-IV and GMM-GLS-IV are similar. Both have the smallest MSE and yield confidence intervals with coverage rates near the nominal level and the shortest length, unlike GMM.
Simulation results with exogenous instruments. RE implies MA(3) errors; general implies ARMA(1,3) errors
| Case | Estimator | MSE | Bias | Variance | Coverage | Length |
|---|---|---|---|---|---|---|
| RE | OLS | 0.19 | 3.43 | 0.17 | 0.93 | 0.16 |
| GMM | 0.21 | 3.66 | 0.08 | 0.77 | 0.11 | |
| GMM-GLS-IV | 0.14 | 2.96 | 0.13 | 0.94 | 0.14 | |
| GLS-IV | 0.15 | 2.99 | 0.12 | 0.93 | 0.13 | |
| General | OLS | 3.66 | 18.05 | 0.37 | 0.19 | 0.24 |
| GMM | 0.79 | 6.98 | 0.31 | 0.75 | 0.22 | |
| GMM-GLS-IV | 0.09 | 2.39 | 0.11 | 0.96 | 0.12 | |
| GLS-IV | 0.10 | 2.42 | 0.11 | 0.95 | 0.12 |
| Case | Estimator | Bias | Variance | Coverage | Length | |
|---|---|---|---|---|---|---|
| 0.19 | 3.43 | 0.17 | 0.93 | 0.16 | ||
| 0.21 | 3.66 | 0.08 | 0.77 | 0.11 | ||
| GMM-GLS-IV | 0.14 | 2.96 | 0.13 | 0.94 | 0.14 | |
| GLS-IV | 0.15 | 2.99 | 0.12 | 0.93 | 0.13 | |
| General | 3.66 | 18.05 | 0.37 | 0.19 | 0.24 | |
| 0.79 | 6.98 | 0.31 | 0.75 | 0.22 | ||
| GMM-GLS-IV | 0.09 | 2.39 | 0.11 | 0.96 | 0.12 | |
| GLS-IV | 0.10 | 2.42 | 0.11 | 0.95 | 0.12 |
Weekly data for US-CAD, US-JP for the period November 2010 to April; 2020 (). For OLS we use HAC standard errors as described in the text
We now consider simulations under the “general case”. We use the DGP (20). However, we consider a departure of the efficient market hypothesis in which the error term is serially correlated at lags . In particular, we assume that follows an process with MA coefficients given by (21) and AR coefficient . The forward rate is generated in the same way as before. We consider the same family of estimators. The results are presented in the second panel of Table 5. In this case, as the error process is correlated beyond lag 3, OLS is no longer consistent. This translates in an important increase in MSE and bias. The confidence intervals are meaningless, in that they have huge size distortions, with a coverage rate smaller than 20%. Since the instruments are exogenous, the GMM estimate remains consistent in this case but is not efficient. The MSE of GMM is eight times that of FGLS. The GMM-GLS-IV estimate provides marked improvements over GMM and achieves important reduction in MSE along with confidence intervals having coverage rates near the nominal level. The finite sample performance of the GLS-IV estimate is similar to GMM-GLS-IV. Both estimates have the smallest MSE and variance with confidence intervals coverage rates near the nominal level and the shortest length.
We next consider the power of the various tests. We consider the DGP (20) for a grid values of around the null hypothesis . In particular, we consider . We set , so that , can be used as instruments. We use the actual observed data for , and we generate according with the DGP (20). The error process is simulated as before: under the “RE case”, follows an process, and under the “general case” follows an process with the same parameter configurations used earlier. In Figures 3 and 4, we plot the empirical rejection frequencies of nominal t-test of , for the same set of estimates considered in the previous section. Figure 3 pertains to the “RE case”, while Figure 4 present results for the “general case”. Note that for the “RE case” with errors case, OLS has good power but is slightly outperformed by the FGLS-based estimates. Remarkably, GMM presents power distortions as it rejects the null hypothesis over 20%. In the “general case” with errors, Figure 4 shows that the power of GMM-GLS-IV and GLS-IV remains the same while that of OLS exhibits huge power distortions skewed at negative values of . The GMM estimate also has large size distortions and skewed power functions.
The multi-line chart displays rejection frequency on the y-axis against beta on the x-axis for 4 estimation methods, G M M-G L S-I V, G L S-I V, O L S, and G M M. All methods show high rejection frequencies close to 1 at beta values away from 0, while rejection frequency drops sharply near beta 0. The G M M-G L S-I V and G L S-I V curves overlap closely, reaching their minimum near 0.08 around beta 0. The O L S curve follows a similar pattern with slightly lower values near the minimum, while the G M M curve maintains comparatively higher rejection frequency around beta 0. The chart highlights strong symmetry around beta 0 and differing sensitivities among the estimation methods.Empirical rejection frequencies of nominal 5% t-Test of . Hansen and Hodrick regression, RE case with MA(3) errors
The multi-line chart displays rejection frequency on the y-axis against beta on the x-axis for 4 estimation methods, G M M-G L S-I V, G L S-I V, O L S, and G M M. All methods show high rejection frequencies close to 1 at beta values away from 0, while rejection frequency drops sharply near beta 0. The G M M-G L S-I V and G L S-I V curves overlap closely, reaching their minimum near 0.08 around beta 0. The O L S curve follows a similar pattern with slightly lower values near the minimum, while the G M M curve maintains comparatively higher rejection frequency around beta 0. The chart highlights strong symmetry around beta 0 and differing sensitivities among the estimation methods.Empirical rejection frequencies of nominal 5% t-Test of . Hansen and Hodrick regression, RE case with MA(3) errors
The line chart presents rejection frequency on the y-axis against beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M estimation methods. The G M M-G L S-I V and G L S-I V curves remain nearly identical, falling sharply to a minimum near beta 0 before rising again towards higher rejection frequencies at positive and negative beta values. The O L S curve differs substantially, increasing steadily from low rejection frequencies at negative beta values to nearly 1 at positive beta values. The G M M curve exhibits a smoother transition with moderate rejection frequencies around beta 0 and gradual increases towards the extremes.Empirical rejection frequencies of nominal 5% t-Test of . Hansen and Hodrick regression, general case with ARMA(1,3) errors
The line chart presents rejection frequency on the y-axis against beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M estimation methods. The G M M-G L S-I V and G L S-I V curves remain nearly identical, falling sharply to a minimum near beta 0 before rising again towards higher rejection frequencies at positive and negative beta values. The O L S curve differs substantially, increasing steadily from low rejection frequencies at negative beta values to nearly 1 at positive beta values. The G M M curve exhibits a smoother transition with moderate rejection frequencies around beta 0 and gradual increases towards the extremes.Empirical rejection frequencies of nominal 5% t-Test of . Hansen and Hodrick regression, general case with ARMA(1,3) errors
In summary, for the case with exogenous instruments, GMM-GLS-IV and GLS-IV clearly have better properties. The performance of OLS is nearly as good in the “RE case” but completely breaks down in the “general case”. Hence, GMM-GLS-IV and GLS-IV are clearly the more robust method of estimation and testing.
6.1.2 Simulations with non-exogenous instruments.
We now consider the case with non-exogenous instruments and first assess the finite sample size of the estimators followed by some power comparisons. The DGP is based on regression of Hansen and Hodrick (1980), considering one-month forward rates ():
where and , . In this case, the set of instruments is simulated so that they are serially correlated and non-exogenous. Accordingly, we set:
for , with independent of . Note that is shock affecting and thus, the instrument is not exogenous whenever . We set and . The process is generated according to the DGP (22) under the null hypothesis . We set so that () can be used as instruments. Under the “RE case”, the error follows an process. Under the “general case”, follows the same process as in Section 6.1. We consider the same set of estimates: OLS + HAC, GMM, GLS-IV and GMM-GLS-IV. The sample size is and 5,000 replications are used.
The simulation results for the “RE case” are presented in the first panel of Table 6. In this case, the OLS estimates are consistent requiring only pre-determined regressors. It is, however, the less efficient estimate and the coverage rates of the associated confidence intervals are below the nominal level. The bad performance of GMM is again corrected when using the GMM-GLS-IV estimate; it achieves the smallest MSE and the coverage rate is better than that of the GLS-IV estimate. The GMM-GLS-IV and GLS-IV estimates have confidence intervals with coverage rates slightly below to the nominal level (92% and 87%, respectively). This results are in line with the Monte Carlo simulation results for the non-exogenous regressors and/or instruments cases in Perron and González-Coya (2022) and Perron and Olivari (2023).
Simulation results with non-exogenous instruments. RE implies MA(3) errors; general implies ARMA(1,3) errors
| Case | Estimator | MSE | Bias | Variance | Coverage | Length |
|---|---|---|---|---|---|---|
| RE | OLS | 0.32 | 4.63 | 0.25 | 0.89 | 0.19 |
| GMM | 0.37 | 4.84 | 0.19 | 0.82 | 0.17 | |
| GMM-GLS-IV | 0.17 | 3.28 | 0.13 | 0.92 | 0.14 | |
| GLS-IV | 0.23 | 3.74 | 0.13 | 0.87 | 0.14 | |
| General | OLS | 3.96 | 18.46 | 0.47 | 0.26 | 0.26 |
| GMM | 1.67 | 10.63 | 0.56 | 0.65 | 0.26 | |
| GMM-GLS-IV | 0.19 | 3.47 | 0.15 | 0.91 | 0.15 | |
| GLS-IV | 0.26 | 3.71 | 0.15 | 0.89 | 0.14 |
| Case | Estimator | Bias | Variance | Coverage | Length | |
|---|---|---|---|---|---|---|
| 0.32 | 4.63 | 0.25 | 0.89 | 0.19 | ||
| 0.37 | 4.84 | 0.19 | 0.82 | 0.17 | ||
| GMM-GLS-IV | 0.17 | 3.28 | 0.13 | 0.92 | 0.14 | |
| GLS-IV | 0.23 | 3.74 | 0.13 | 0.87 | 0.14 | |
| General | 3.96 | 18.46 | 0.47 | 0.26 | 0.26 | |
| 1.67 | 10.63 | 0.56 | 0.65 | 0.26 | ||
| GMM-GLS-IV | 0.19 | 3.47 | 0.15 | 0.91 | 0.15 | |
| GLS-IV | 0.26 | 3.71 | 0.15 | 0.89 | 0.14 |
Weekly data for US-CAD, US-JP for the period November 2010 to April; 2020 (). For OLS we use HAC standard errors as described in the text
The simulation results for the “general case” with errors are presented in the second panel of Table 6. Since the serial correlation in the errors extends beyond lag 3, OLS and GMM are no longer consistent. This is reflected in large MSE, bias and variance. The size distortions are exacerbated with coverage rates of the confidence intervals below 65% (26% for OLS). The finite sample performance of the GLS-IV estimate is in line with the fact that it is consistent. Surprisingly, despite the fact that the first step GMM estimates of the GMM-GLS-IV procedure are not consistent, the resulting GMM-GLS-IV estimate has smaller MSE than the GLS-IV estimate. These results suggest that the FGLS-based procedures are very robust to the first-stage estimates of the autocorrelation coefficients. The MSE of GLS-IV is almost 12 times smaller than the MSE of OLS, and it achieves with a variance that is on average five times smaller than the variance of the other estimates. The coverage rates of the confidence intervals of the GMM-GLS-IV and GLS-IV estimates are near the nominal level with the smallest length.
We now consider the power analysis. We consider the DGP (22) for a grid values of around the null hypothesis . In particular, we consider , . We set , so that , can be used as instruments. We generate according with the DGP (22). Under the “RE case”, follows an process, and under the “general case” with an process. In Figures 5 and 6, we plot the empirical rejection frequencies of nominal t-test of , for the same set of estimates considered in the previous section. Figure 5 pertains to the “RE case”, while Figure 6 pertains to the “general case with errors.
The chart illustrates rejection frequency on the y-axis versus beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M methods. The G M M-G L S-I V and G L S-I V curves overlap almost completely, decreasing from near 1 to their minimum around beta 0 before increasing symmetrically again. The O L S curve shows lower rejection frequencies near positive beta values compared with the other methods, while the G M M curve lies between the O L S and G L S-I V curves across most beta values. The overall pattern demonstrates a pronounced V-shaped behaviour centred around beta 0.Empirical rejection frequencies of nominal 5% t-Test of . Non-exogenous IVs, RE case with MA(3) errors
The chart illustrates rejection frequency on the y-axis versus beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M methods. The G M M-G L S-I V and G L S-I V curves overlap almost completely, decreasing from near 1 to their minimum around beta 0 before increasing symmetrically again. The O L S curve shows lower rejection frequencies near positive beta values compared with the other methods, while the G M M curve lies between the O L S and G L S-I V curves across most beta values. The overall pattern demonstrates a pronounced V-shaped behaviour centred around beta 0.Empirical rejection frequencies of nominal 5% t-Test of . Non-exogenous IVs, RE case with MA(3) errors
The line chart shows rejection frequency on the y-axis plotted against beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M methods. The G M M-G L S-I V and G L S-I V curves are nearly identical, exhibiting a deep minimum near beta 0 and rejection frequencies approaching 1 at larger absolute beta values. The O L S curve rises steadily from low rejection frequencies at negative beta values to high rejection frequencies at positive beta values. The G M M curve demonstrates moderate rejection frequencies across the range, increasing gradually from negative to positive beta values. The comparison highlights differing rejection behaviours and asymmetry among the estimation approaches.Empirical rejection frequencies of nominal 5% t-Test of . Non-exogenous IVs, General case with ARMA(1,3) errors
The line chart shows rejection frequency on the y-axis plotted against beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M methods. The G M M-G L S-I V and G L S-I V curves are nearly identical, exhibiting a deep minimum near beta 0 and rejection frequencies approaching 1 at larger absolute beta values. The O L S curve rises steadily from low rejection frequencies at negative beta values to high rejection frequencies at positive beta values. The G M M curve demonstrates moderate rejection frequencies across the range, increasing gradually from negative to positive beta values. The comparison highlights differing rejection behaviours and asymmetry among the estimation approaches.Empirical rejection frequencies of nominal 5% t-Test of . Non-exogenous IVs, General case with ARMA(1,3) errors
From the results in Figure 5 for errors, all the tests have similar power functions, though somewhat lower for OLS and GMM when is positive. The FGLS-based procedures exhibit the same power function, with a rejection frequency of the null hypothesis that is slightly higher than the 5% nominal level. For GMM the null rejection frequency is higher than 15%. The results in Figure 6 pertaining to the errors case show that the power functions of the FGLS-based procedures exhibit the same behavior as in the errors case. On the other hand, OLS and GMM are now subject to important size distortions. The power function of those estimates is biased toward negative values of . Note that OLS rejects the null hypothesis almost 10% of the times when , while it rejects with a frequency higher than 75% when . Hence, the only reliable test are those obtained using the FGLS-based procedures.
Remark 3. Three consistent messages emerge across all Monte Carlo designs reported in Tables 1–6, spanning both the Fama and Hansen–Hodrick specifications, with and without exogenous instruments, and for one- and three-month forward rates. First, when the error process is MA() – the rational expectations case – all estimators are consistent, but FGLS-based procedures achieve substantially smaller MSE and variance than OLS, with gains ranging from a factor of 3–5 depending on the specification. OLS with HAC standard errors delivers confidence intervals near the nominal coverage level in this case, but at the cost of considerably wider intervals and correspondingly lower power (Figures 1, 3 and 5). Second, when the error process exhibits serial correlation beyond lag – the general case with ARMA() errors – OLS becomes inconsistent, its MSE increases by an order of magnitude and its confidence intervals suffer severe size distortions with coverage rates falling below 20% in some designs. The FGLS-based procedures remain consistent and efficient, with coverage rates near the nominal level and MSE largely unchanged relative to the rational expectations case. Third, among the IV-based procedures required for specifications with lagged dependent variables, GMM-GLS-IV and GLS-IV deliver similar finite-sample performance, with both dominating GMM in terms of MSE, coverage and power. The overall practical implication is clear: the FGLS-based procedures are more efficient and at worst marginally more complex than OLS when the rational expectations hypothesis holds, and decisively superior when it does not, making them the dominant choice for h -step-ahead forecasting regressions in applied work.
6.2 Empirical results
We now report estimates of the model (19) using weekly spot exchange rates for the UK pound (US-UK), Canadian dollar (US-CAD) and Japanese yen (US-JP). We use the complete US-JP sample that spans from June 1995 to January 2023, with 1,442 observations. We consider the OLS + HAC, GMM, GMM-GLS-IV and GLS-IV estimates. For the FGLS-based methods we set . The estimation results for one-month forward rates () are presented in Table 7, while those for three-month forward rates () are in Table 8. As documented in Section 7, the outcome of the Cumby and Huizinga (1992) test (CH-test) for autocorrelation at lags show a rejection in all cases for the residuals of the Hansen and Hodrick (1980) regression (19), which opens the possibility of correlation between the regressors/instruments and the past errors.
Hansen and Hodrick (1980) one-month forward model estimation results (SE in parentheses)
| FX | OLS | GMM | GMM-GLS-IV | GLS-IV | OLS | GMM | GMM-GLS-IV | GLS-IV |
|---|---|---|---|---|---|---|---|---|
| US-UK | −0.000 (0.001) | 0.001 (0.002) | 0.000 (0.001) | 0.000 (0.001) | −0.087 (0.040) | 0.044 (3.875) | 1.436 (1.439) | 0.392 (0.278) |
| US-CAD | −0.000 (0.001) | 0.000 (0.001) | −0.000 (0.001) | −0.000 (0.001) | −0.035 (0.051) | 0.247 (1.012) | 0.040 (0.631) | 1.223 (0.959) |
| US-JP | 0.003 (0.002) | 0.004 (0.022) | 0.001 (0.001) | 0.007 (0.005) | 0.120 (0.054) | 0.024 (6.181) | −0.039 (0.598) | −0.534 (1.060) |
| US-UK | 0.143 (0.059) | 0.029 (1.650) | −0.442 (0.435) | −0.081 (0.135) | 0.016 (0.045) | 0.025 (0.599) | −0.112 (0.133) | −0.037 (0.039) |
| US-CAD | 0.012 (0.046) | −0.096 (0.280) | 0.053 (0.162) | −0.455 (0.358) | 0.031 (0.037) | 0.050 (0.074) | 0.023 (0.024) | 0.034 (0.031) |
| US-JP | −0.056 (0.063) | −0.042 (1.262) | −0.045 (0.092) | 0.050 (0.187) | 0.015 (0.058) | 0.013 (0.421) | −0.027 (0.044) | 0.028 (0.045) |
| | ||||||||
|---|---|---|---|---|---|---|---|---|
| FX | GMM-GLS-IV | GLS-IV | GMM-GLS-IV | GLS-IV | ||||
| US-UK | −0.000 (0.001) | 0.001 (0.002) | 0.000 (0.001) | 0.000 (0.001) | −0.087 (0.040) | 0.044 (3.875) | 1.436 (1.439) | 0.392 (0.278) |
| US-CAD | −0.000 (0.001) | 0.000 (0.001) | −0.000 (0.001) | −0.000 (0.001) | −0.035 (0.051) | 0.247 (1.012) | 0.040 (0.631) | 1.223 (0.959) |
| US-JP | 0.003 (0.002) | 0.004 (0.022) | 0.001 (0.001) | 0.007 (0.005) | 0.120 (0.054) | 0.024 (6.181) | −0.039 (0.598) | −0.534 (1.060) |
| US-UK | 0.143 (0.059) | 0.029 (1.650) | −0.442 (0.435) | −0.081 (0.135) | 0.016 (0.045) | 0.025 (0.599) | −0.112 (0.133) | −0.037 (0.039) |
| US-CAD | 0.012 (0.046) | −0.096 (0.280) | 0.053 (0.162) | −0.455 (0.358) | 0.031 (0.037) | 0.050 (0.074) | 0.023 (0.024) | 0.034 (0.031) |
| US-JP | −0.056 (0.063) | −0.042 (1.262) | −0.045 (0.092) | 0.050 (0.187) | 0.015 (0.058) | 0.013 (0.421) | −0.027 (0.044) | 0.028 (0.045) |
For US-UK, is the coefficient for US-CAD, and is the coefficient for US-JP; for US-CAD, is the coefficient for US-UK and is the coefficient for US-JP; for US-JP, is the coefficient for US-UK and is the coefficient for US-CAD
Hansen and Hodrick (1980) three-month forward model estimation results (SE in parentheses)
| FX | OLS | GMM | GMM-GLS-IV | GLS-IV | OLS | GMM | GMM-GLS-IV | GLS-IV |
|---|---|---|---|---|---|---|---|---|
| US-UK | −0.001 (0.004) | 0.000 (0.004) | −0.000 (0.000) | −0.001 (0.001) | 0.091 (0.077) | 0.937 (2.486) | 0.536 (0.424) | 0.376 (0.212) |
| US-CAD | 0.000 (0.003) | 0.000 (0.004) | −0.000 (0.001) | −0.000 (0.001) | 0.005 (0.071) | 0.103 (1.379) | 1.063 (0.604) | 0.594 (0.206) |
| US-JP | −0.001 (0.004) | −0.002 (0.006) | 0.000 (0.001) | 0.002 (0.004) | 0.104 (0.062) | 0.112 (0.703) | 1.090 (0.919) | 2.349 (1.372) |
| US-UK | 0.055 (0.090) | −0.378 (1.028) | −0.102 (0.134) | −0.104 (0.123) | −0.026 (0.059) | −0.146 (0.408) | −0.067 (0.043) | −0.054 (0.027) |
| US-CAD | 0.026 (0.070) | −0.044 (0.420) | −0.290 (0.158) | −0.239 (0.097) | 0.097 (0.046) | 0.088 (0.130) | 0.017 (0.037) | 0.085 (0.023) |
| US-JP | −0.051 (0.106) | −0.038 (0.254) | −0.193 (0.124) | −0.342 (0.199) | 0.064 (0.108) | 0.111 (0.219) | −0.074 (0.059) | −0.062 (0.128) |
| FX | GMM-GLS-IV | GLS-IV | GMM-GLS-IV | GLS-IV | ||||
|---|---|---|---|---|---|---|---|---|
| US-UK | −0.001 (0.004) | 0.000 (0.004) | −0.000 (0.000) | −0.001 (0.001) | 0.091 (0.077) | 0.937 (2.486) | 0.536 (0.424) | 0.376 (0.212) |
| US-CAD | 0.000 (0.003) | 0.000 (0.004) | −0.000 (0.001) | −0.000 (0.001) | 0.005 (0.071) | 0.103 (1.379) | 1.063 (0.604) | 0.594 (0.206) |
| US-JP | −0.001 (0.004) | −0.002 (0.006) | 0.000 (0.001) | 0.002 (0.004) | 0.104 (0.062) | 0.112 (0.703) | 1.090 (0.919) | 2.349 (1.372) |
| US-UK | 0.055 (0.090) | −0.378 (1.028) | −0.102 (0.134) | −0.104 (0.123) | −0.026 (0.059) | −0.146 (0.408) | −0.067 (0.043) | −0.054 (0.027) |
| US-CAD | 0.026 (0.070) | −0.044 (0.420) | −0.290 (0.158) | −0.239 (0.097) | 0.097 (0.046) | 0.088 (0.130) | 0.017 (0.037) | 0.085 (0.023) |
| US-JP | −0.051 (0.106) | −0.038 (0.254) | −0.193 (0.124) | −0.342 (0.199) | 0.064 (0.108) | 0.111 (0.219) | −0.074 (0.059) | −0.062 (0.128) |
For US-UK, is the coefficient for US-CAD, and is the coefficient for US-JP; for US-CAD, is the coefficient for US-UK and is the coefficient for US-JP; for US-JP, is the coefficient for US-UK and is the coefficient for US-CAD
In subsection 6.1.2, we provided evidence that the GMM-GLS-IV and GLS-IV estimates are consistent requiring only pre-determined instruments . We shall thus focus on the FGLS-based estimates. First note that both estimates are similar across all currencies and forecast horizons. For one- and three-month forward rates, the FGLS-based estimates cannot reject the null hypothesis for any of the currencies. In contrast, OLS rejects for US-JP, for one- and three-month forward rates. There is no consensus in the literature about the rejection of this null hypothesis. For three-month forward rates (), Hansen and Hodrick (1980) using OLS rejects the null hypothesis for US-CAD and two other currencies (Deutsche mark and Swiss franc) for data between 1975 and 1979.
Note the relatively high standard errors of the GMM estimates for all the currencies. This suggests that the instruments may be non-exogenous. On the other hand, despite the fact that the first stage GMM estimate in the GMM-GLS-IV procedure is very noisy, the resulting GMM-GLS-IV estimate is much more efficient and close to the GLS-IV estimate. It can be argued that while the GLS-IV estimate is valid with only pre-determined instruments, it might still be subject to a “weak instrument” problem; see, e.g. Andrews et al., 2019. We provide statistical evidence to argue that this is not the case. In particular, we consider the test of Staiger and Stock (1997) for weak instruments and the Wu–Hausman exogeneity test (Hausman, 1978 and Wu, 1973). The detailed implementation for the GLS-IV estimate are outlined in Appendix A.4.1 and A.4.2. Note that standard weak instruments tests are valid for GLS-IV as the 2SLS regression (12) has uncorrelated errors. Table 9 reports the weak instruments test, the test statistic and the exogeneity test , see (A.5), together with the corresponding p-values. The null hypothesis for weak instruments is rejected for all currencies and at the level of significance. The Wu-Hausman rejects the null hypothesis of exogeneity for US-UK and US-JP at least at the level of significance for in all cases. Overall, these results indicate that we can be confident about the estimates and tests obtained using the GLS-IV procedure when applied to the Hansen and Hodrick (1980) regression (19).
Weak instruments (first row) and Wu–Hausman (second row) tests for GLS-IV
| Forward rates | FX | Statistic | p-value | ||
|---|---|---|---|---|---|
| 1-month | US-UK | 72 | 1309 | 13.02 | 2E-16 |
| 1 | 1379 | 7.91 | 0.00499 | ||
| US-CAD | 78 | 1299 | 7.785 | 2E-16 | |
| 1 | 1375 | 4.022 | 0.044 | ||
| US-JP | 48 | 1349 | 17.90 | 2E-16 | |
| 1 | 1395 | 30.18 | 4.67E-08 | ||
| 3-month | US-UK | 75 | 1288 | 22.75 | 2E-16 |
| 1 | 1361 | 56.27 | 1.13E-13 | ||
| US-CAD | 75 | 1288 | 24.689 | 2E-16 | |
| 1 | 1361 | 10.67 | 0.001 | ||
| US-JP | 39 | 1348 | 18.745 | 2E-16 | |
| 1 | 1385 | 9.066 | 0.00265 |
| Forward rates | FX | | | Statistic | p-value |
|---|---|---|---|---|---|
| 1-month | US-UK | 72 | 1309 | 13.02 | 2E-16 |
| 1 | 1379 | 7.91 | 0.00499 | ||
| US-CAD | 78 | 1299 | 7.785 | 2E-16 | |
| 1 | 1375 | 4.022 | 0.044 | ||
| US-JP | 48 | 1349 | 17.90 | 2E-16 | |
| 1 | 1395 | 30.18 | 4.67E-08 | ||
| 3-month | US-UK | 75 | 1288 | 22.75 | 2E-16 |
| 1 | 1361 | 56.27 | 1.13E-13 | ||
| US-CAD | 75 | 1288 | 24.689 | 2E-16 | |
| 1 | 1361 | 10.67 | 0.001 | ||
| US-JP | 39 | 1348 | 18.745 | 2E-16 | |
| 1 | 1385 | 9.066 | 0.00265 |
7. Uncovering the OLS bias
As shown in subsection 2.1, if the error term in regression (1) is serially correlated at lags , the OLS estimator is no longer consistent. In this section we briefly present the results from applying the Cumby and Huizinga (1992) test (CH-test) for autocorrelation at lags to both the Fama (1984) regression (13) and the Hansen and Hodrick (1980) regression (19). This test is well suited for our purpose since the null hypothesis is that the error process is a moving average of known order against the general alternative that the autocorrelations are nonzero at lags greater than q. The CH-Test is a Wald test of the null hypothesis that the regression error is uncorrelated with itself at lags through . A general formulation for two-stage least squares and two-step two-stage least squares is presented in Cumby and Huizinga (1992). We require the errors to be unconditionally homoskedastic. We refer to the paper by Cumby and Huizinga (1992) and Appendix A.4.3 for the details about the implementation of the test, which require the estimate of several quantities. We simply note that for the Fama regression (13) we use FGLS residuals:
where and , whereas for the Hansen and Hodrick (1980) regression (19) we use GLS-IV residuals:
where and for . In Table 10, we provide the results of the CH-Test statistics, labelled , of the null hypothesis that the regression error in the Fama regression (13) with 3-month forward rates is uncorrelated with itself at lags to with and . For both sample periods, we obtain a rejection of the null hypothesis for every s and conclude that the error term in regression (13) is serially correlated at lags . Table 11 provide similar results for the Hansen and Hodrick (1980) regression (19). The specifications are similar except that we set . Again, for both sample periods, we have statistical evidence to reject the null hypothesis for every s. Hence, again here the error term in regression (19) is also serially correlated at lags .
Cumby and Huizinga (1992) test of the null hypothesis that the Fama regression error is uncorrelated with itself at lags to ,
| Period | Lags | US-UK | US-CA | US-JP |
|---|---|---|---|---|
| 10/1984–01/2023 | 12 | 196.499 | 321.282 | 223.436 |
| 15 | 70.894 | 48.850 | 61.324 | |
| 20 | 57.016 | 47.526 | 44.952 | |
| 10/1989–04/2021 | 12 | 152.768 | 280.233 | 642.996 |
| 15 | 34.919 | 48.315 | 29.323 | |
| 20 | 29.680 | 75.373 | 54.568 |
| Period | Lags | US-UK | US-CA | US-JP |
|---|---|---|---|---|
| 10/1984–01/2023 | 12 | 196.499 | 321.282 | 223.436 |
| 15 | 70.894 | 48.850 | 61.324 | |
| 20 | 57.016 | 47.526 | 44.952 | |
| 10/1989–04/2021 | 12 | 152.768 | 280.233 | 642.996 |
| 15 | 34.919 | 48.315 | 29.323 | |
| 20 | 29.680 | 75.373 | 54.568 |
Cumby and Huizinga (1992) test of the null hypothesis that the Hansen–Hodrick regression error is uncorrelated with itself at lags to ,
| Period | Lags | US-UK | US-CA | US-JP |
|---|---|---|---|---|
| 10/1984–01/2023 | 5 | 1087.218 | 203.5084 | 646.771 |
| 10 | 1134.685 | 219.5272 | 624.1312 | |
| 50 | 525.2874 | 330.3852 | 1115.21 | |
| 10/1989–04/2021 | 5 | 48811.75 | 43522.55 | 61825.54 |
| 10 | 114112.9 | 424805.4 | 65447.43 | |
| 50 | 93889.12 | 377855.8 | 37567.25 |
| Period | Lags | US-UK | US-CA | US-JP |
|---|---|---|---|---|
| 10/1984–01/2023 | 5 | 1087.218 | 203.5084 | 646.771 |
| 10 | 1134.685 | 219.5272 | 624.1312 | |
| 50 | 525.2874 | 330.3852 | 1115.21 | |
| 10/1989–04/2021 | 5 | 48811.75 | 43522.55 | 61825.54 |
| 10 | 114112.9 | 424805.4 | 65447.43 | |
| 50 | 93889.12 | 377855.8 | 37567.25 |
8. Practical recommendations
Our FGLS framework applies to any setting in which a researcher estimates the parameters of a h-step-ahead linear forecasting equation using data sampled at a finer frequency than the forecast horizon. This structure arises naturally in a wide range of economic and financial applications beyond the UIP tests studied here. Examples include: (i) forecasting inflation or output growth at quarterly horizons using monthly data, where overlapping forecast errors induce MA() serial correlation under rational expectations; (ii) predictive regressions for equity returns at horizons exceeding the sampling frequency, as in the stock return predictability literature; e.g. Campbell and Shiller (1988) and Stambaugh (1999); (iii) survey-based forecast evaluation, where professional forecasters issue multi-period-ahead predictions at regular intervals and the econometrician seeks to test forecast rationality or efficiency; and (iv) macro-finance term structure models in which bond risk premia are regressed on predictive variables at horizons longer than the observation frequency. In each of these settings, the key question is whether the forecast error is serially correlated only up to lag – as implied by rational expectations – or beyond. If the latter, OLS is inconsistent when the regressors are not strictly exogenous, and the FGLS procedures developed here provide a consistent and efficient alternative.
We offer the following decision framework to guide practitioners in choosing among the estimators considered in this paper.
Step 1: Determine whether the regressors include lagged dependent variables. If the regression includes only contemporaneous or exogenous regressors, as in the Fama specification (13), the basic FGLS procedure of Section 3 applies directly. The Durbin regression is estimated via OLS, the autoregressive coefficients are used to quasi-difference the data, and the FGLS estimate is obtained from the quasi-differenced regression (7). No instrumental variables are required. This is the simplest and most efficient procedure when it applies.
Step 2: If lagged dependent variables are present, assess instrument exogeneity. When the regression includes lagged dependent variables, as in the Hansen-Hodrick specification (18), the choice between GMM-GLS-IV and GLS-IV depends on the plausible exogeneity of the available instruments. If the practitioner has strong reasons to believe the instruments are exogenous (i.e. uncorrelated with the error process at all leads and lags), GMM-GLS-IV is appropriate. It uses a first-stage GMM estimate to obtain consistent residuals, from which the autoregressive filtering coefficients are estimated. The advantage is that the first-stage GMM estimate and the associated residuals are consistent under exogeneity, yielding a clean identification of the autoregressive parameters.
However, if exogeneity of the instruments cannot be credibly maintained – and in many economic applications it cannot – GLS-IV should be preferred. GLS-IV requires only that the instruments to be pre-determined (uncorrelated with contemporaneous and future errors, but potentially correlated with past errors), a strictly weaker condition. Pre-determination is the natural assumption in most time series settings where instruments are lagged values of observable variables. As shown in our simulation experiments (Section 6.1.2), GLS-IV remains consistent and achieves competitive efficiency even with non-exogenous instruments, whereas GMM and GMM-GLS-IV can exhibit substantial bias and size distortions in this case. A notable finding from our simulations is that GMM-GLS-IV can perform well even when its first-stage GMM estimate is inconsistent (Table 6), suggesting that the FGLS filtering step is robust to moderate contamination of the first-stage residuals. Nevertheless, in the absence of a formal test confirming instrument exogeneity, we recommend GLS-IV as the default choice for its broader robustness guarantees.
Step 3: Diagnostic testing. Regardless of which estimator is selected, we recommend two diagnostic checks. First, the Cumby–Huizinga test (Section 7) should be applied to the estimated residuals to assess whether serial correlation extends beyond lag . If the null hypothesis of MA() errors is not rejected, OLS with HAC standard errors remains consistent and the gains from FGLS are purely in efficiency. If the null is rejected, as we find in all our empirical specifications, the FGLS-based procedures are necessary for consistency. Second, when using GLS-IV, the Staiger–Stock weak instruments test and the Wu–Hausman exogeneity test (Section A.4) should be reported to assess instrument strength and to provide evidence on whether the IV correction is empirically relevant.
Step 4: Lag order selection. The maximum lag order for the BIC selection of should satisfy the rate condition as . In practice, we recommend as a starting point. For our sample sizes of approximately 1,500–2,000 weekly observations, this yields between 11 and 13, though we use the more generous values of or 40 to allow the BIC sufficient flexibility. Practitioners working with shorter samples (e.g. monthly data with ) should use smaller values of to avoid overfitting. Following Ng and Perron (2005), the BIC comparison across different values of p should use the same effective number of observations to ensure a proper comparison.
9. Concluding remarks
We re-examined the statistical evidence about the hypothesis of UIP, which is a joint hypothesis of efficiency in the forward foreign exchange markets and rational expectations. Testing rationality hypothesis and exchange market efficiency is embedded in the general problem of estimating the parameters of a h-step-ahead linear forecasting equation. Under the null hypothesis, the forecast errors are serially correlated up to lags and OLS is consistent. However, if the errors are serially correlated beyond lags , OLS is no longer consistent. This observation motivates using FGLS-based methods that are robust to the structure of the error process. We apply the FGLS procedure developed in Perron and González-Coya (2022) and we extend it to a setting with lagged dependent variables included as regressors. The resulting instrumental variables-based procedure, GLS-IV, is consistent requiring only pre-determined IVs. Using these FGLS methods, we study the main two UIP specifications in the literature: the Fama (1984) and Hansen and Hodrick (1980) regressions. We provide novel insights about the forward premium anomaly. Applying the consistent FGLS method we show that the estimates of for three currencies are always non negative at the 1% significance level. A result that is contrary to the general finding that the OLS estimates are negative. Hence, the so-called “forward discount anomaly” is not as severe as previously thought. We also show statistically significant discrepancies between the OLS and GLS-IV estimates in the Hansen and Hodrick (1980) regression. We rationalize these discrepancies by showing that the regression residuals are in fact serially correlated beyond lags k, and thus OLS is not consistent, while the FGLS methods remain consistent. This point to the usefulness of adopting our more robust FGLS procedure. Not only is consistent and efficient under a wider range of contexts but, as we have shown, can deliver estimates that are different and point to a different assessment of the empirical facts.
The traditional interpretation of the forward discount anomaly rests on OLS estimates of in the Fama regression that are negative, with an average around across 75 published estimates Froot (1990). As Fama (1984) showed, this implies that the variance of the risk premium must exceed the variance of expected depreciation, a condition that has proven notoriously difficult to reconcile with standard asset pricing models. Much of the subsequent theoretical literature on habit formation, long-run risks and rare disasters has been motivated, at least in part, by the need to generate risk premia large enough to explain these extreme negative values. Our FGLS estimates reframe the problem. With positive but generally below unity, the implied risk premium is still present but far more modest. The gap between and one is consistent with a small, positive covariance between the risk premium and the forward discount; this is in line in the sign and magnitude that calibrated consumption-based models can plausibly deliver without requiring extreme parameter values; see, e.g. Lustig and Verdelhan (2007) and Bansal and Shaliastovich (2013). In other words, the “puzzle” that much of the risk premium literature has sought to resolve was largely an artifact of inconsistent OLS estimation inflating the apparent size of the premium. The existing theoretical toolkit may already be adequate to explain the risk premia implied by our consistent estimates, rendering the anomaly considerably less anomalous than previously thought.
References
Appendix. A.1 Proof of Remark 2
Assume that has a permanent and a transitory component, , where the permanent component is defined by , where , and are assumed to be independent. Then we can write:
Hence:
Thus, we can write the forecast error as , where . Note that and can be allowed to be correlated, so that is invertible in general.
A.2 Efficient estimate of autoregressive coefficients for GLS-IV
Suppose that . The efficient linear combination is obtained when minimizes . Hence, the minimization problem is , subject to The first-order conditions of this problem are:
where is the Lagrangian multiplier. The solution is:
A.3 Simulation design
We simulate an process with parameters calibrated to replicate the observed autocorrelation function up to lag of the FGLS residuals of regression (13) using US-UK data for the Fama regression (Section 5) and using GLS-IV residuals of regression (19) for the Hansen and Hodrick (1980) regression (Section 6). We obtain the parameters by solving the non-linear system of equations given by the autocorrelation functions (ACF) for . Let be an process with , then the variance of is as follows:
The autocovariance function of is as follows:
Thus, the autocorrelation function of is as follows:
Using the observed first q autocorrelations of the residuals of the regression, () we define a system of q non-linear equations with q unknown variables :
A.4 Diagnostic tests for GLS-IV
Rewrite the model using the quasi-differenced variables , and as:
where (A.1) is the structural equation of interest, is a vector and is a vector with the lagged dependent variable (the only endogenous variable in the model). (A.2) is the reduced form equation for Y, W is the matrix of exogenous regressors with row t, and (i.e. W includes a constant term). Z is the matrix of quasi-differenced instruments with row t, and . and V are, respectively, a vector and a matrix of error terms. Note that the quasi-differenced regression (A.1) has serially uncorrelated errors, with covariance matrix . Assume that and . Let , it is assumed throughout that .
A.4.1 Staiger and Stock (1997) weak instruments test
We are interested in testing in the regression (A.2). shall be modeled as local to zero, so that the F statistic is . Staiger and Stock (1997) make the assumption that where c is a fixed vector. Before proceeding we provide some additional definitions and notation. Let , partitioned so that , and . Also let . Let and where R is a general matrix with , and let “” denote the residuals from the projection on W, so , , etc. Let and and let denote the k-dimensional identity matrix. Staiger and Stock (1997) assume that the following limits hold jointly: 1) ; 2) ; 3):
where , with . Define , where :
and . The random variable is distributed , where is the matrix with and , where is partitioned conformably with . Finally, let:
and:
The 2SLS estimate of is:
By standard projection arguments, the 2SLS estimate of is as follows:
The Wald statistic testing is , where , with . Staiger and Stock (1997) show that the limit distribution of is defined in (A.3). As we just have one endogenous variable, , the F statistic converges to a non-central with noncentrality parameter . In the general case, with more than one endogenous variable, is the matrix of noncentrality parameters of the limiting noncentral Wishart random variable .
A.4.2 Wu–Hausman test of exogeneity
The Wu–Hausman (WH) test (see Hausman, 1978 and Wu, 1973) examines the null hypothesis that Y is exogenous (i.e. ) by checking for a statistically significant difference between the OLS and 2SLS estimates of . The test statistic is as follows:
With:
Its limit distribution is as follows:
where and . Under the null hypothesis , simplifies to:
where with and and are independent. Note that since , applying critical values to results in asymptotically conservative tests. However, as noted by Staiger and Stock (1997), a size adjustment of is infeasible because the distribution depends on .
A.4.3 Testing for serial correlation at lags q > k−1
We describe in some details the test of Cumby and Huizinga (1992) for autocorrelation structure of the OLS residuals (CH-Test). This test is perfectly suited for our purpose as it allows to have under the null hypothesis a regression error process with a moving average of known order against the general alternative that the autocorrelations of the regression error are nonzero at lags greater than q. The CH-Test is a Wald test of the null hypothesis that the regression error is uncorrelated with itself at lags through . Consider the general formulation of the CH-Test for an OLS regression. Here we present the Cumby and Huizinga (1992) test for autocorrelation structure of the OLS residuals of equation (1) with no instrumental variables. A general formulation for two-stage least squares and two-step two-stage least squares is presented in Cumby and Huizinga (1992). The model is as follows:
where is a vector of the n scalar predetermined regressors. The regression errors, are assumed to be serially correlated up to a known lag and their autocorrelations at all lags greater than q are required to be zero under the null hypothesis. We require the errors to be unconditionally homoskedastic. The CH-Test statistic is as follows:
where are consistent estimates of . Here, r is a vector :
is the asymptotic covariance matrix of the estimator :
with . Let D be the matrix . B is the matrix with i, jth element:
Let for for , the ijth element of the matrix be given by , and the ijth element of the matrix C be given by . To consistently estimate the test statistic , we need consistent estimate of the errors . For the Fama regression (13), we use FGLS residuals:
where and . Whereas for the Hansen and Hodrick (1980) regression (19), we use GLS-IV residuals:
where and for . We estimate as discussed in subsection 4.1.1.

