Purpose

This paper aims to re-examine the uncovered interest parity (UIP) hypothesis, which posits efficiency in forward foreign exchange and rational expectations. Testing these assumptions involves estimating parameters in a k-step-ahead forecasting model, where forecast errors are expected to be serially correlated up to lags k 1. When errors are correlated beyond these lags, OLS is no longer consistent unless the regressors are exogenous.

Design/methodology/approach

The authors extend the FGLS procedure developed in Perron and González-Coya (2022) to a setting in which lagged dependent variables are included as regressors. The authors thus provide a consistent and efficient framework to estimate the parameters of a general k-step-ahead linear forecasting equation. Following the work of Perron and Olivari (2023), the authors introduce an instrumental variable (IV)-based approach for this problem that requires pre-determined but not necessarily exogenous IVs for consistency.

Findings

The authors apply the authors’ FGLS procedures to the analysis of the two main specifications to test the UIP. Contrary to most empirical results available in the literature, in particular those based on some OLS regression or GMM, the authors’ robust and efficient procedure cannot reject the null hypothesis that the UIP holds.

Originality/value

Overall, this study’s results can be viewed as overturning the so-called forward discount anomaly. The methods proposed can also be applied to a wide variety of contexts.

In this paper, we re-examine the hypothesis of uncovered interest parity (UIP), which in its basic form implies that the (nominal) expected return to speculation in the forward foreign exchange market conditional on available information should be zero. This is an “efficient-markets hypothesis” (EMH) for foreign exchange markets: if all available information is used rationally by risk-neutral agents in determining the spot and forward exchange rates, then the expected rate of return to speculation will be zero and the foreign exchange market is said to be efficient. This is a joint hypothesis since it includes the assumption of rational expectations (REH) and the assumption that the risk premium for the forward rate is zero. In fact, rejection of the UIP hypothesis does not immediately translates into a rejection of the efficiency of the foreign exchange market, which could be due economic agents being risk averse. Still testing whether the UIP hypothesis holds has been and continue to be a topic of considerable interest from both theoretical and empirical perspectives.

Testing the rationality hypothesis and exchange market efficiency is embedded in the general problem of estimating the parameters of a k -step-ahead linear forecasting equation. When the sampling interval is finer than the interval over which forecasts are made (in this case the maturity time of the forward exchanges rates), the forecast error is serially correlated. As noted by Hansen and Hodrick (1980), under rational expectations (REH) the forecast error is serially correlated up to lag k1 and OLS remains consistent but appropriate modifications in the estimation of the asymptotic covariance matrix are needed. However, if the REH is rejected and the forecast error is serially correlated beyond lag k1, OLS is no longer consistent when the regressors are not exogenous; see Perron and González-Coya (2022). Contrary to what is asserted in Hansen and Hodrick (1980), GLS is consistent when the regressors are pre-determined provided the roots of the MA polynomial are inside the unit circle, i.e. MA process is invertible, as shown in Perron and González-Coya (2022). Moreover, GLS remains consistent when the forecast error follows a linear invertible process.

The first contribution of this paper is to provide a consistent and efficient framework to estimate and perform tests of the parameters of a k -step-ahead linear forecasting equation that remains valid whether the REH holds or not. We apply the FGLS procedure developed in Perron and González-Coya (2022) which is consistent using non-exogenous regressors, provided the errors follows an stationary invertible linear process. The second contribution is to extend their FGLS procedure to cases with lagged dependent variables included as regressors. Following the work of Perron and Olivari (2023), we introduce an instrumental variable (IV)-based approach for this problem that requires pre-determined but not necessarily exogenous IVs for consistency. The third contribution is to apply our FGLS procedures to the two main regressions suggested in the literature to test the UIP. We use 30 years of data for three currencies, and we reconsider the framework and regressions used by Fama (1984) and Hansen and Hodrick (1980). We provide extensive simulation experiments to assess the finite sample performance of our FGLS procedure relative to OLS. We show that FGLS achieves important reductions in mean-squared error (MSE) and allow tests with much greater power.

The Fama regression assesses whether the current forward-spot differential, ft,hst, is a good predictor of the future change in the spot rate, st+hst. Most results available in the literature suggest a negative estimate of the relevant parameter, which instead should take value one if the UIP holds. This is often referred to as the “forward discount anomaly”, which refers to the widespread empirical finding that the returns on nominal exchange rates is negatively correlated with the lagged forward premium. It implies an appreciating currency for the high interest rate country. It is an “anomaly” as rational expectations would imply the opposite; if all currencies are equally risky, investors would demand higher interest rates on currencies expected to fall in value. Our results, in contrast, indicate positive values, sometimes not significantly different from one. This finding suggests that the “forward discount anomaly” might be a consequence of OLS providing an inconsistent estimate and our FGLS procedure being consistent and efficient under a broader range of possible scenarios.

The regression adopted by Hansen and Hodrick (1980) is to test whether past values of the forward-spot differential help predict the current value ft,hst, conditioning on some covariates involving the past forward-spot differentials from some other countries. Under the UIP and EMH, there should be no predictive power as all information contained in the information set at time t should already have been accounted for by the market in setting the forward rates. Here, contrary to most empirical results available in the literature, in particular those based on some OLS regression or GMM, our robust and efficient procedure cannot reject the null hypothesis that the UIP holds. Hence, overall, our results can be viewed as overturning the so-called “forward discount anomaly.”

The main methodological contribution of this paper is that it extends Perron and González-Coya (2022)FGLS procedure for the case in which lagged dependent variables are included as regressors (subsection 3.1). In subsection 4.2, we discuss an application of Perron and Olivari (2023) FGLS-IV method when the instrumental variable set includes lagged pre-determined regressors as instruments. These contributions, provide a consistent and efficient framework to estimate and perform tests of the parameters of a k-step-ahead linear forecasting equation that remains valid whether the REH holds or not. Empirically, the most significant contribution is that by applying the feasible GLS and GLS-IV methodologies described above to the estimation of the Fama and Hansen and Hodrick (1980) regressions, respectively, we obtain results that widely differ from the consensus in the literature. For the Fama regression, most results available in the literature suggest a negative estimate of the relevant parameter, often referred to as the “forward discount anomaly.” Our results in contrast, indicate positive values, sometimes not significantly different from one. This finding suggests that the “forward discount anomaly” might be a consequence of OLS providing an inconsistent estimate and our FGLS procedure being consistent and efficient under a broader range of possible scenarios. We also show statistically significant discrepancies between the OLS and GLS-IV estimates in the Hansen and Hodrick (1980) regression.

The remainder of this paper is as follows. Section 2 describes the formulations of the UIP hypothesis and reviews previous econometric tests. Section 3 describes the FGLS procedures and Section 4 introduces two instrumental variable (IV) based approaches for a setting in which lagged dependent variables are included as regressors. Section 5 studies the Fama (1984) regression and Section 6 is focused on the Hansen and Hodrick (1980) regression. For both specifications, we provide extensive Monte Carlo experiments to assess the finite sample performance of our FGLS procedures relative to OLS. We also provide empirical results for the sample period January 1993 to January 2023. In Section 7, we present a test for the hypothesis that the OLS residuals exhibits serial correlation of order greater than k1 and discuss its empirical implications to understand the various conflicting results. In Section 8, we provide practical recommendations and brief concluding remarks are presented in Section 9. A supplement provide additional details.

Let st=ln(St) and ft,k=ln(Ft,k), where St and Ft,k are the levels of the spot exchange rate and the k-period forward exchange rate at time t. With it, the domestic nominal interest rate and it the corresponding foreign interest rate, the theory of UIP implies that E(st+kst|Φt)=(itit). Hence, UIP requires the twin assumptions of rational expectations and a constant or zero risk premium. Given the no arbitrage condition, the covered interest parity (CIP) condition implies that (itit)=(ft,kst) holds as an identity. Hence, the UIP condition is also frequently expressed as E(st+kst|Φt)=(ft,kst). Since st+kft,k is an approximate measure of the rate of return to speculation, we can express the efficient-markets hypothesis as ft,k=E(st+k|Φt). This implies forecast errors st+kft,k uncorrelated with information available at time t, Φt.

As in Hansen and Hodrick (1980), we consider the general problem of estimating the parameters of a k-step-ahead linear forecasting equation, E(yt+k|Φt)=xtβ. Then:

(1)

where rational expectations impose a specific structure on the forecast error ut+k=yt+kE(yt+k|Φt). Due to the k1 period overlap in the sequential k-step-ahead forecasts, ut+k has an MA(k1) representation. Thus:

and the OLS estimates of α,β are consistent since the regressors are pre-determined via the rational expectations hypothesis. We shall label this as the “RE case”. Note, however, that we could well be faced with a model of the form:

where ηt is some serially correlated process involving innovations dated before period t; e.g. an AR(1) process of the form ηt=ρηt1+et for some sequence of i.i.d. innovations (e1,,eT). In this case, if the regressors are non-exogenous with respect to past values of ηt, OLS is no longer consistent since E(xtηt)0. We shall label this as the “general case”, as it encompasses the “RE case”. This motivates the necessity of an estimate that is consistent under the both the “RE case” and the “general case”, i.e. allowing the errors to be serially correlated beyond lag k1. Contrary to what is asserted in Hansen and Hodrick (1980), GLS is consistent in both cases, provided the errors (ut+k+ηt) can be represented as some invertible linear process, as shown in Perron and González-Coya (2022).

Example 1. A k-step ahead forecast error process can be serially correlated beyond lagsk1under, e.g. Adaptive Expectations (AE) (seeMuth, 1960). An exponentially weighted moving average forecast arises from the following model of expectations adapting to changing conditions. Under AE, it is assumed that the forecast is changed from one period to the next by an amount proportional to the latest observed error:

As shown inMuth (1960), the solution of this difference equation is an exponentially weighted forecast yte=κi=1(1κ)i1yti. The forecast error is thus:

To characterize the process of the forecast error ut, we shall impose a functional form on yt. It is standard in the AE literature, to assume that yt has a permanent and a transitory component, yt=y¯t+ωt,where the permanent component is defined by y¯t=y¯t1+εt=i=1tεiwithεti.i.d(0,σε2), ωti.i.d(0,σω2) and εt,ωt independent. In this case, the forecast error ut follows an AR(1) process. The details of this derivation are spelled in the  Appendix A.1.

Several tests of the UIP hypothesis have been proposed in the literature. Bilson (1981) and Fama (1984) analyzed regression (1) with yt+k=st+kst and xt=ft,kst with k=1. In this setting a test of UIP is that H0:α=0 and β=1. It has been noted by Fama (1984) and many subsequent studies that the estimated slope coefficient β is frequently negative. This is known as the Forward Premium Anomaly: the country with the higher rate of interest has an appreciating currency rather than a depreciating currency; a violation of UIP. Hansen and Hodrick (1980) estimate the model (1) with yt+k=st+kft,k, the forecast error, and uses lagged dependent variables as regressors; xt=(yt,yt1)) with k=13. The test is H0:α=0 and β1=β2=0. A simplified version of this regression, with xt=yt and test H0:α=0 and β1=0, has been studied by Baillie et al. (2023). Hansen and Hodrick (1980) also proposed a test that includes the lagged forecast errors of four other currencies:

(2)

where stiftk,ki is the forecast error for country i and stjftk,kj is the forecast error for country j ≠ i. We focus on regression (2) as Hansen and Hodrick (1980) concludes that the multicountry test appears to be more powerful.

We consider the linear regression:

(3)

where the error term follows a short-memory linear process:

(4)

where εti.i.d.(0,σ2). The polynomial θ(L)=j=0θjL j is assumed to satisfy θ0=1 (a normalization), j=0j|θj|< so that the process is short-memory and θ(L) is invertible, i.e. we can write θ(L)1ut+k=εt+k. We use the FGLS procedure developed in Perron and González-Coya (2022) to estimate regression (3). The idea is to consistently approximate ut+k using an autoregression of order kT:

with kT and kT3/T0 as T, to ensure consistent estimates; see Berk (1974). Replacing equation (3) and re-arranging terms we have the following:

(5)

with δj=βρj for j=1,,kT. Equation (5) is often called the Durbin regression (see Durbin, 1970). The order of the autoregression, kT, is determined via the minimization of the BIC suggested by Schwarz (1978) for kT[0,kmax] where kmax is such that kmax3/T0 as T. We use the method suggested by Ng and Perron (2005) to ensure a proper comparison across models with different values of kT, i.e. using the same effective number of observations. We estimate the Durbin equation (5) via OLS with kt=kT. Using the OLS estimates of ρj, ρ^jD, we construct the quasi-differenced variables:

(6)

The FGLS estimate of β is the OLS estimate of the quasi-differenced regression:

(7)

The resulting estimate β^FGLS will be consistent provided the regressors xt are pre-determined with respect to values εt prior to period t+k; see Perron and González-Coya (2022).

Remark 1. The consistency of our FGLS and GLS-IV estimators is unaffected by conditional heteroskedasticity. The key requirement for consistency is that the error process ut+h admits a stationary invertible linear representation, so that the autoregressive filtering recovers serially uncorrelated errors and the quasi-differenced regressors remain pre-determined. Our variance estimators and the analogous IV-based expressions to be discussed below assume unconditional homoskedasticity of the quasi-differenced errors. Under conditional heteroskedasticity, these estimators remain consistent for the unconditional variance of the coefficient estimates.

Consider regression (3) with xt=(yth,wt) for some h>0, where wt is a vector of n+1 pre-determined regressors that includes a constant term. For simplicity we omit the constant term without loss of generality. Write the regression as:

(8)

If ut is autocorrelated beyond lags ih1, OLS applied to regression (8) is not consistent as yth and ut are not independent. Wallis (1967) and Malinvaud (1966) studied regression (8) with h=1 where the error terms follows an AR(1) process. Expression for the asymptotic bias of the OLS estimates of α and β are given by Malinvaud (1966) and Griliches (1961). We can write the Durbin regression as follows:

(9)

where γh=ρh+β, γj+h=ρj+hδj+h, for j=1,,kTh; δj+h=ρj+hβ, for j=kTh+1,,kT; and ψij=αiρj for j=1,,kT, i1,,n. For the case h=kT=1 and wt a scalar (i.e. n=1), Wallis (1967) and Malinvaud (1966) (p. 469) propose to estimate regression (9) using OLS and then estimate ρj using ψj/α. For the general case k>1 and kT>h, the same approach can be applied. Note that we need kTh, otherwise the estimates are not consistent. But this condition will be satisfied, at least in large samples, when using the BIC to select the lag order. However, note that when n>1, ρj cannot be uniquely identified. We shall propose two estimation methods; one based on the Durbin regression (9) and one based on a first-stage instrumental variable (IV) estimate. These follow similar steps as in the feasible GLS procedure discussed above.

Remark 2. Estimating the autoregressive coefficients, ρjforj=1,,kT, from regression (9) using the Wallis (1967) and Malinvaud (1966) method does not allow us to uniquely identify ρj in the general case with k>1, kT>kt and n>1. However, we can obtain efficient estimates ρ˜j, j=1,,kT by using a convex combination of the estimates ρ˜ij=ψ˜ij/α˜i (j=1,,kT) for i1,,n. For ease of the exposition, suppose that n=2. Then, we can construct efficient estimates ρ˜j=λρ˜1j+(1λ)ρ˜2j, where the optimal λ that minimize Var(ρ˜j) is (the details are in the  Appendix A.2):

In practice, we can estimate Var(ρ˜ij) using a first order Taylor expansion:

The first order Taylor expansions of Cov(ρ˜1j,ρ˜2j) are cumbersome. Note that if the instruments wtj are independent of each other, Cov(ρ˜1j,ρ˜2j) will be arbitrarily small in large samples. Hence, we set Cov(ρ˜1j,ρ˜2j)=0.

While the adaptive expectations framework provides an analytically tractable illustration of how forecast errors can be serially correlated beyond lag h1, we emphasize that it is one instance of a broader class of departures from the rational expectations paradigm that generate such excess autocorrelation. Several economically motivated models of expectations formation share this property. Under sticky information models Mankiw and Reis (2002), agents update their information sets infrequently, so that aggregate expectations adjust sluggishly to new information. This generates forecast errors that inherit the persistence of the underlying fundamentals, producing autocorrelation well beyond the h1 lags implied by overlapping forecasts alone. Similarly, noisy information models in the tradition of Sims (2003) and Woodford (2003), in which agents observe fundamentals with noise and must solve a signal extraction problem, produce forecast errors whose serial correlation structure reflects the dynamics of the signal-to-noise ratio rather than the forecast horizon alone. Models with heterogeneous expectations – where agents use different forecasting rules and the population weights evolve over time – also generate aggregate forecast errors with rich autocorrelation structures; see, e.g. De Grauwe and Grimaldi (2006). Finally, behavioral models featuring extrapolative or momentum-based expectations Barberis et al. (1998) can produce persistent forecast errors when agents systematically over- or under-react to recent exchange rate movements. The key insight is that our FGLS procedure does not require the econometrician to specify or identify which of these mechanisms is operative. It requires only that the error process be representable as a stationary invertible linear process. This agnosticism is a practical advantage: the procedure delivers consistent and efficient estimates regardless of whether the excess serial correlation originates from adaptive learning, rational inattention, heterogeneous beliefs or any other mechanism that generates a departure from the pure MA(h1) structure.

An alternative to estimate regression (8) with n1 is to use an instrumental variable procedure. If the regressors wt are exogenous, then wth are valid instruments for yth and the two-stage least squares (2SLS) estimates will be consistent. Liviatan (1963) propose to use wth as instruments for yth and wt as an instrument for itself, so that the instrument set is Zt={wth,wt}. We can potentially select a larger set of instrumental variables for yth; e.g. lags or order ih of wt. In this case zt={wt,wti,p>ih}, for some p>h. Note that p can be adaptively selected using appropriate tests for over-identifying restrictions (see Small, 2007). However, if the regressors wt are not exogenous and ut is autocorrelated beyond lags ih1, wth are no longer valid instruments. Hence, the IV estimate of regression (8) using the instrument set zt will not be consistent. This motivates the use of the GLS-IV procedure suggested by Perron and Olivari (2023). The idea is simple: first transform the model to have serially uncorrelated errors and then estimate the transformed model via IV using the set of transformed instruments. The resulting estimate will be consistent as the instrument set and regressors are pre-determined in the transformed model.

We first discuss two methods that are valid with exogenous instruments if ut is autocorrelated beyond lags h1. One is the widely used so-called “optimal GMM” procedure. The other is akin to the GLS procedure discussed above but with the first-step using the GMM estimate to obtain estimate of the residuals and construct the autoregressive filtering.

4.1.1 GMM.

Using the set of n(ph) exogenous instruments zt, then, as shown in Hansen (1982), the best estimator of β=(β,α) based on the instruments and the moment condition E(Zu)=0 is as follows:

(10)

where the n(ph)×n(ph) matrix Ω is given by Ω=limTT1E[ZuuZ]. We can write Ω=s=Rv(s), where Rv(s)=E[Zh(t)utZh(ts)uts] and Zh(t)={zt,,zth}. Then Ω can thus be consistently estimated by Ω^=s=TTλ(s,m)R^v(s), where R^v(s)=T1t=1T|s|vtvt+|s|, with vt=Zh(t)ut, λ(s,m) is some kernel or weight function and m is the bandwidth; see, e.g. Andrews (1991).

4.1.2 GMM-GLS-IV.

We also consider a FGLS method that does not rely on the Durbin regression (9). Instead, it uses a first-stage GMM estimate to obtain a consistent estimate of the residuals, u˜t. We can thus identify the autoregressive parameters when the instruments are exogenous. The GMM-GLS-IV procedure to estimate regression (8) with exogenous instruments wt is the following: 1) obtain the GMM estimator of regression (8), given by (10) using the set of instruments zt={wt,wth,wthj,j=1,,h}. Compute the residuals, u˜t=ytα˜ivwtβ˜ivyth; 2) select the order kT of the autoregression:

(11)

via the minimization of the BIC for kT[0,kmax] where kmax is such that kmax3/T0 as T; 3) estimate the autoregression (11) with kT=kT to obtain consistent estimates ρ˜j (j=1,,k); 4) use ρ˜j (j=1,kT) to construct the quasi-differenced variables yt=(ytj=1kTρ˜jytj), yth=(ythj=1kTρ˜jythj) and wt=(wtj=1kTρ˜jwtj); 5) the GMM-GLS-IV estimate of β is the IV estimate of the quasi-differenced regression:

(12)

using the set of quasi-differenced instruments zt={wt,wth}.

We next describe the GLS-IV procedure to estimate regression (8), which is valid with non-exogenous instruments, provided they are pre-determined. The steps are the following: 1) Select the order of the Durbin regression (9), kT via the minimization of BIC for kT[0,kmax] where kmax is such that kmax3/T0 as T; 2) Estimate the Durbin regression (9) with the selected value kT using OLS. The estimates the autoregressive coefficients, ρ˜j, j=1,,kT are obtained using the efficient method described in Remark 2; 3) Use ρ˜j, j=1,kT to construct the quasi-differenced variables yt, ytk and wt, as in Step 4 for GMM-GLS-IV; 4) The GLS-IV estimate of β is the IV estimate of the quasi-differenced regression (12) using the set of quasi-differenced instruments zt={wt,wth}.

Given that the OLS estimates from the Durbin regression (9) are consistent and that the transformed model (12) has serially uncorrelated errors, the GLS-IV estimates are consistent under the stated conditions. To the best of our knowledge, there is no other consistent estimation method requiring only pre-determined regressors for the general linear regression (8) without restricting the error process ut. Maximum likelihood and non-linear least squares require an a priori known error process. In the simulation experiments reported in subsection 6.1.2, we show that the fact that ρ˜j is not uniquely identified when we have more than one regressor (n>1), does not affect the efficiency of GLS-IV.

The Fama (1984) regression is as follows:

(13)

with h=4 for 1-month forward rates and h=12 for 3-month forward rates, when using weekly data. Estimates of (13) tell us whether the current forward-spot differential, ft,hst, has power to predict the future change in the spot rate, st+hst. Evidence that β is significantly different from zero means that the forward rate observed at t has information about the spot rate to be observed at t+h. Under the efficient-market hypothesis, we have H0:α=0,β=1. Under the null hypothesis, the log of the forward rate provides an unbiased forecast of the log of the future spot exchange rate. Derivations from β=1 are sometimes interpreted as a measure of the variation of the premium in the forward rate.

Let β¯ be the OLS estimate of β in regression (13). If the estimator is consistent, we have as follows:

(14)

If expectations are rational, then st+1st=Et(st+1)st+εt+1, where εt+1=st+1Et(st+1) is the forecast error. In this case, Cov(ft,hst,st+1st)=Cov(ft,hst,Et(st+1)st). The foreign exchange risk premium when expectations are rational is defined as rptre=ft,hEt(st+1). Under risk neutrality, expected profits from forward market speculation would be zero as agents would drive ft,h into equality with Et(st+1). Write Et(st+1)st=ft,hstrptre and replace into equation (14), so that plimT(β¯)=1βrp, where:

A negative estimate of β in regression (13) is a robust finding in the literature (see Engel, 1996). This is known as the “forward discount anomaly”; it is a widespread empirical finding that the returns on nominal exchange rates appear to be negatively correlated with the lagged forward premium. Bilson (1981) and Fama (1984) provide evidence that the estimates of β are less than zero. Many subsequent studies have confirmed that finding, for dollar exchange rates and a large number of exchange rates and time periods (see, for example, Bekaert and Hodrick, 1993; Backus et al., 1993; Hai et al., 1997). Froot (1990) notes that the average value of β¯ over 75 published estimates is −0.88. Only a few of the estimates are greater than zero, and none is greater than 1. The forward discount anomaly implies an appreciating currency for the high interest rate country. It is an “anomaly” as rational expectations would imply the opposite; if all currencies are equally risky, investors would demand higher interest rates on currencies expected to fall in value. The survey by Engel (1996) focuses on the possibility that β¯rp ≠ 0 among the possible explanations for finding β¯<0. Other possible interpretations are that the forward rate is a biased predictor of the future spot rate, and/or that it is evidence of a time-varying risk premium. In this paper, we provide a novel insight on this empirical regularity. We argue that the OLS and GMM estimates of β are not consistent, as the error term in regression (13) is serially correlated beyond lags k1. We find that the FGLS estimates, which are consistent under a broader range of conditions, are significantly non negative but in general smaller than 1.

Daily data were obtained for the spot exchange rates for the UK pound (US-UK), Canadian dollar (US-CAD) and Japanese yen (US-JP) as well as the one- and three-month forward exchange rates data for the three currencies. As in Hansen and Hodrick (1980), the data were sampled to form a weekly series constructed by taking observation on Tuesday of each week. If no Tuesday observation was available, we used the Wednesday observations. The source of the forward exchange data for US-UK and US-CAD is Barclays Bank PLC; the source for US-JP is the Bank of Tokyo Mitsubishi. For all the data sets, we use all the information available until January 17, 2023 but the starting date of the time series differ: a) US-UK: starting date of October 11, 1983, 2,050 observations; b) US-CAD: starting date of December 14, 1984, 1,989 observations; c) US-JP: starting date of September 1, 1993, 1,535 observations.

In this section, we provide simulation results related to the Fama (1984) regression (13) under the “RE case” or efficient market hypothesis (EMH), H0:α=0,β=1. We consider h=4 for one-month forward rates and h=12 for three-month forward rate. As discussed in Section 2, the EMH implies that ut+h has an MA(h1) representation. We simulate an MA(h1) process with parameters calibrated to replicate the observed autocorrelation function up to lag h1 of the FGLS residuals of equation (13) using US-UK data. The details are presented in the  Appendix A.3. For h=4 we have an MA(3) representation with:

(15)

For h=12 we have an MA(11) representation with coefficients:

(16)

We simulate an error process based on equation (15), i.e. ut+h=θ(L)εt, where εti.i.d.N(0,σε2) and σε2 is estimated using the residuals of an initial Fama FGLS regression (13) for US-UK. The data generating process (DGP) uses xt=ft,hst observed in the data for US-UK, and the stated simulated error process. We artificially generate yt+4 to satisfy the null hypothesis H0:β=1. Thus, the DGP is as follows:

(17)

We also consider a departure from the “RE case” with errors ut+h following the “general case” with serial correlation at lags ih. We assume that ut+h follows an ARMA(1,h) process with MA coefficients given by (15) and (16) depending on h=4 and h=12, respectively. The AR coefficient is set to ρ=0.6. We consider two sampling periods; the complete sample from October 11, 1983 to January 17, 2023 with 2,050 observations and the one spanning November 1. 1989 to April 1, 2021, as considered in Baillie et al. (2023). We perform 5,000 replications. We consider the following estimators of α,β, from regression (13): a) The OLS estimate with HAC standard errors based on the quadratic spectral window suggested by Andrews (1991) with automatic bandwidth selection using an ARMA(1,1) approximation; b) The Durbin estimate based on the regression (5) with yt=st+hst and xt={1,ft,hst} and kT selected using the BIC. For the complete sample, we consider kmax=40 and for the sub-sample we set kmax=30; c) The FGLS estimate based on regression (7) with yt=st+hst, xt={1,ft,hst} and kT selected using the BIC. The same values of kmax are used. We estimate the sample variance of the FGLS estimator using Var(γ^FGLS)=(XX)1σ^FGLS2, where X is the T×2 matrix of quasi-differenced regressors (including a constant term) and σ^FGLS2 is the sample variance of the FGLS residuals, u^FGLS=ytγ^FGLSxt. The confidence intervals at the 95% nominal level for the jth coefficient are obtained using γ^j±z0.975Var(γ^j,FGLS)1/2, where z0.975 is the 0.975 quantile of the normal distribution.

The simulation results are presented in Table 1 for h=4 and Table 2 for h=12. In line with the theory, the mean squared error (MSE) of OLS is small when the error term follows an MA(h) process and deteriorates when the error is serially correlated at lags ih. Clearly, FGLS outperforms OLS and Durbin in all cases. Even under the “RE case”, the MSE of OLS is on average 3.4 times larger than of FGLS, while the MSE of Durbin is close to that of OLS case. When the errors follow an ARMA(1,h) process, OLS is no longer efficient, and its MSE is on average 27 times that of FGLS. As expected, the Durbin estimate remains consistent in this case, but its MSE is on average three times that of FGLS. FGLS also has the smallest variance. The variance of Durbin is on average three times the variance of FGLS for h=4 and two times the variance of FGLS for h=12. The coverage rates of the confidence interval for FGLS are near the nominal 90% and have shortest lengths. OLS with HAC standard errors exhibits substantial size distortions with ARMA(1,h) errors, especially when h=12 case. The OLS based HAC standard errors provides confidence intervals close to the nominal level in some cases, at the expense of a very large variance. For the ARMA(1,h) case with k=4(12) the variance of OLS is 30 (25) times the variance of FGLS. This results are in line with those from Perron and González-Coya (2022).

Table 1.

Simulation results under EMH, one-month forward rate. RE implies MA(3) errors; general implies ARMA(1,3) errors

PeriodCaseEstimator MSEBiasVarianceCoverageLength
11/83–01/23REOLS0.81520.73040.75980.943.3952
FGLS0.18010.33880.19190.961.7167
Durbin0.49500.55530.50570.952.7888
GeneralOLS4.92391.79744.47380.938.2112
FGLS0.14450.30680.15250.961.5311
Durbin0.49270.00350.49550.952.7607
01/89–04/21REOLS1.00630.79291.00190.943.8879
FGLS0.31080.44220.33470.952.2667
Durbin0.98380.78831.04710.964.0135
GeneralOLS6.22711.97305.83640.929.3351
FGLS0.27360.42120.29260.962.1200
Durbin0.99240.79291.05020.954.0196
Note(s):

We use US-UK observed data xt=ft,kst, T1=2,050, T2=1,600. For OLS we use HAC standard errors as described in the text

Table 2.

Simulation results under EMH, three-month forward rate. RE implies MA(11) errors; general implies ARMA(1,11) errors

Period CaseEstimator MSEBiasVarianceCoverageLength
11/83–01/23REOLS1.68981.05411.50360.924.7404
FGLS0.70980.67610.72610.953.3387
Durbin1.39410.93581.41480.954.6647
GeneralOLS18.57443.497215.32600.9115.0322
FGLS0.60190.62830.63320.963.1202
Durbin1.39640.93701.41930.954.6720
01/89–04/21REOLS1.88951.08691.65580.914.9454
FGLS0.99950.78170.98340.943.8838
Durbin1.98431.15022.09730.965.6801
GeneralOLS20.61803.593316.21220.8815.3295
FGLS0.88250.74510.90270.953.7256
Durbin1.97821.14672.09950.965.6832
Note(s):

We use US-UK observed data xt=ft,kst, T1=2,050, T2=1,600. For OLS we use HAC standard errors as described in the text

For the simulations related to power, we use US-UK observed data xt=ft,hst for the period November 1, 1989 to April 1, 2021 with one-month forward rates, h=4. The DGP is as follows:

(18)

where β{0,0.25,0.5,0.75,1,1.25,1.5,1.75,2}. The null value is β=1. The error process ut+4 is simulated as before; under the “RE case” ut+4 follows an MA(3) process, and under the “general case” ut+4 follows an ARMA(1,3) process. Figures 1 and 2 present plots of the empirical rejection frequencies for t-test of H0:β=1 with nominal size 0.05 for the OLS estimates with HAC-ARMA(1,1) standard errors, the FGLS estimates based on regression (7) and the Durbin estimates based on the regression (5). For the latter two, kT is selected via BIC with kmax=30. Figure 1 pertains to “RE case” with the errors an MA(3) process, while Figure 2 pertains to the “general case” with errors following an ARMA(1,3) process.

Figure 1.
A line chart compares rejection frequency versus beta for F G L S, O L S, and Durbin methods.The line chart presents rejection frequency on the y-axis against beta on the x-axis for 3 statistical methods, F G L S, O L S, and Durbin. The F G L S curve begins at the highest rejection frequency near 0.58 at beta 0, decreases sharply to approximately 0.05 around beta 1, and then rises again to about 0.41 at beta 2. The O L S and Durbin curves show lower rejection frequencies throughout, both declining towards beta 1 before increasing gradually at higher beta values. The chart demonstrates a U-shaped trend for all methods, with F G L S exhibiting the greatest variation and highest rejection frequency overall.

Empirical Rejection Frequencies Of Nominal 5% t-Test of H0:β=1. Fama regression, RE case with MA(3) errors

Figure 1.
A line chart compares rejection frequency versus beta for F G L S, O L S, and Durbin methods.The line chart presents rejection frequency on the y-axis against beta on the x-axis for 3 statistical methods, F G L S, O L S, and Durbin. The F G L S curve begins at the highest rejection frequency near 0.58 at beta 0, decreases sharply to approximately 0.05 around beta 1, and then rises again to about 0.41 at beta 2. The O L S and Durbin curves show lower rejection frequencies throughout, both declining towards beta 1 before increasing gradually at higher beta values. The chart demonstrates a U-shaped trend for all methods, with F G L S exhibiting the greatest variation and highest rejection frequency overall.

Empirical Rejection Frequencies Of Nominal 5% t-Test of H0:β=1. Fama regression, RE case with MA(3) errors

Close modal
Figure 2.
A line chart compares rejection frequency versus beta for F G L S, O L S, and Durbin methods.The line chart shows rejection frequency on the y-axis plotted against beta on the x-axis for F G L S, O L S, and Durbin statistical methods. The F G L S curve starts above 0.60 at beta 0, decreases steeply to its minimum near beta 1, and then rises again towards beta 2. The O L S curve remains comparatively stable between 0.07 and 0.11 across most beta values, while the Durbin curve declines from approximately 0.26 to around 0.05 before increasing moderately at higher beta values. The results indicate that F G L S is more sensitive to changes in beta than the other methods.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=1. Fama regression, General case with ARMA(1,3) errors

Figure 2.
A line chart compares rejection frequency versus beta for F G L S, O L S, and Durbin methods.The line chart shows rejection frequency on the y-axis plotted against beta on the x-axis for F G L S, O L S, and Durbin statistical methods. The F G L S curve starts above 0.60 at beta 0, decreases steeply to its minimum near beta 1, and then rises again towards beta 2. The O L S curve remains comparatively stable between 0.07 and 0.11 across most beta values, while the Durbin curve declines from approximately 0.26 to around 0.05 before increasing moderately at higher beta values. The results indicate that F G L S is more sensitive to changes in beta than the other methods.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=1. Fama regression, General case with ARMA(1,3) errors

Close modal

Note that in Figure 1, “RE case” with MA(3) errors, OLS and Durbin has very small rejection frequencies even when β is far from the null value 1, even though they are consistent, which can be attributed to their lack of efficiency in finite samples. As shown in Figure 2, the power functions are similar in the case with ARMA(1,3) errors, with the exception that the power function of OLS decreases and flattens with an almost constant 10% rejection frequency for all values, despite having a more liberal size.

We present the estimation results of regression (13) using the three estimates considered before: OLS, Durbin estimates based on the regression (5) and FGLS based on regression (7). We use data for three currencies, US-UK, US-CAD and US-JP and we consider two sampling periods; the first one is the complete sample and the second one spans from November 1, 1989 to April 1, 2021, the sampling period considered in Baillie et al. (2023). For the Durbin and FGLS estimates we set kmax=40 for the full sample and kmax=30 for the sub-sample period. Tables 3 and 4 present the estimation results using one-month forward rates (h=4) and the three-month forward rates (h=12), respectively. In line with the empirical regularity in the literature, the OLS estimates of β are not significantly positive in all cases. In contrast, the Durbin and FGLS estimates are positive in most cases and significantly non negative in some. In some cases, such as the US-JP exchange rate for three-month forward rates, we observe a negative significant OLS estimate and a positive significant FGLS estimate. Recall that the Durbin and FGLS estimates are consistent even when the error follows a general linear process. Section 7 present the outcome of the Cumby and Huizinga (1992) test (CH-test) for autocorrelation at lags q>h1. The outcome is a rejection in all cases. Hence, our finding suggests that the forward discount anomaly might be a consequence of OLS providing an inconsistent estimate of β caused by non-exogenous regressors. We thus interpret the large differences between the OLS and the FGLS estimates as evidence of OLS being inconsistent.

Table 3.

Fama (1984) model estimation results, one-month forward rate (SE in parentheses)

 α β
FX PeriodOLSDurbinFGLSOLSDurbinFGLS
US-UK11/1983–01/2023−0.00167 (0.001)−0.00031 (4.1E-04)−0.00006 (3.3E-04)−1.08291 (0.316)−0.09390 (0.380)0.08457 (0.191)
10/1989–04/2021−0.00241 (0.002)−0.00067 (0.001)−0.00027 (0.001)−1.32248 (0.359)0.44659 (0.781)−0.28232 (0.464)
US-CA12/1984–01/2023−0.00015 (0.001)−0.00000 (2.5E-04)0.00004 (2.1E-04)−0.24314 (0.333)0.14125 (0.376)0.20060 (0.212)
10/1989–04/2021−0.00005 (0.002)0.00000 (0.001)0.00005 (0.001)0.04296 (0.384)0.49843 (0.652)0.19765 (0.375)
US-JP09/1993–01/2023−0.00165 (0.002)−0.00018 (0.001)0.00005 (3.3E-04)−0.08744 (0.110)0.26966 (0.145)0.17845 (0.081)
09/1993–04/2021−0.00101 (0.003)−0.00027 (0.001)0.00011 (0.001)−0.13273 (0.110)0.14601 (0.282)0.31225 (0.152)
Table 4.

Fama (1984) model estimation results, three-month forward rate (SE in parentheses)

 α β
FX PeriodOLSDurbinFGLSOLSDurbinFGLS
US-UK11/1983–01/2023−0.00439 (0.004)−0.00035 (4.1E-04)0.00001 (3.3E-04)−0.98614 (0.218)0.64187 (0.382)0.29696 (0.278)
10/1989–04/2021−0.00011 (0.004)−0.00011 (0.001)0.00018 (0.001)0.34508 (0.238)1.25038 (0.474)1.15847 (0.332)
US-CA12/1984–01/2023−0.00036 (0.003)0.00004 (2.5E-04)0.00006 (2.1E-04)−0.21093 (0.218)0.51211 (0.343)0.31376 (0.261)
10/1989–04/2021−0.00040 (0.003)0.00002 (0.001)−0.00002 (0.001)0.19149 (0.252)0.55918 (0.399)0.31524 (0.301)
US-JP09/1993–01/2023−0.00355 (0.005)−0.00017 (0.001)−0.00015 (3.3E-04)−0.10423 (0.138)0.45446 (0.152)0.30124 (0.099)
09/1993–04/2021−0.00158 (0.005)−0.00003 (0.001)−0.00002 (0.001)−0.26883 (0.139)0.49440 (0.157)0.32649 (0.101)

The economic intuition behind our findings can be summarized as follows. The forward discount anomaly – the robust negative estimate of β in the Fama regression – has traditionally been interpreted as reflecting either a time-varying risk premium that covaries negatively with the forward discount, a failure of rational expectations or some combination of the two. Our results point to a fundamentally different explanation: the anomaly is primarily a statistical artifact arising from the inconsistency of OLS when the forecast error is serially correlated beyond lag h1. To see why, note that in the Fama regression the regressor xt=ft,hst is not strictly exogenous with respect to the full error process ut+h. When ut+h exhibits serial correlation at lags lh – a feature we confirm empirically via the Cumby–Huizinga test (Tables 10 and 11) – past realizations of u are correlated with current values of xt through their joint dependence on fundamentals that drive both the forward discount and subsequent spot rate movements. This induces a negative bias in the OLS estimate of β, mechanically pushing it below zero. The FGLS procedure eliminates this bias by filtering the regression to produce serially uncorrelated errors, thereby restoring the orthogonality between the transformed regressors and transformed errors that is required for consistency. The resulting positive estimates of β are thus not an anomalous new finding; they are what one should expect once the estimation procedure is robust to the actual structure of the error process. In other words, the forward rate does contain information about future spot rates broadly consistent with UIP, but this relationship is obscured in standard OLS regressions by an endogeneity bias induced by excess serial correlation in the forecast errors. The decades-long literature documenting negative β estimates may thus have been attributing to economic forces – large and volatile risk premia, systematic expectational failures – what is more parsimoniously explained as a consequence of inconsistent estimation.

Our FGLS estimates of β are positive across all currencies and forward rate maturities, but in most specifications they remain below unity. This pattern warrants discussion in terms of the risk premium. Recall that under rational expectations, plim(β^)=1Δ, where Δ captures the covariance between the risk premium rpt and the forward discount ft,hst, scaled by the variance of the forward discount. A consistent estimate of β that is positive but less than one implies Δ>0, which is consistent with a time-varying risk premium that covaries positively with the forward discount, which is the sign predicted by standard asset pricing theory. Under the consumption-based CAPM or its Epstein–Zin generalization, currencies with higher interest rates are expected to depreciate, but investors require a positive premium for bearing the associated consumption risk, so the forward rate exceeds the expected future spot rate by a modest margin. Our estimates suggest that this risk premium component, while present, is economically small. To illustrate, consider the US-JP estimates for three-month forward rates (Table 4): the FGLS estimate of β=0.30 implies Δ ≈ 0.70, whereas the OLS estimate of β=0.10 implies Δ ≈ 1.10. The latter would require the variance of the risk premium to exceed the variance of the expected depreciation – a condition that Fama (1984) himself noted was difficult to reconcile with plausible models of risk compensation. Our FGLS estimates, by contrast, imply a more moderate risk premium whose magnitude is broadly consistent with calibrated models of international asset pricing; see, e.g. Lustig and Verdelhan (2007) and Bansal and Shaliastovich (2013). We note, however, that our framework does not allow us to separately identify the risk premium and expectational components. The departure of β^FGLS from unity could reflect a genuine, economically meaningful risk premium, a residual (but small) departure from rational expectations or both. What our results do establish is that the magnitude of this departure is far smaller than previously documented, and in particular that the implausibly large negative estimates of β from OLS are attributable to inconsistent estimation rather than to economic fundamentals.

We now turn to an estimation problem that shares some of the main features, though with added complexities. Our aim is to efficiently estimate Hansen and Hodrick (1980) regression:

(19)

where yt+h=st+hift,hi and wtj=stjfth,hj for j ≠ i. Note that if αj is significantly different from zero, then wthj is correlated with the regression variable yt and can be potentially used as an instrument. If the lagged forecast error of country j ≠ i, wthj, is exogenous, i.e. uncorrelated with the residuals us for all t and s then GMM and GMM-GLS-IV are consistent. If wthj is only pre-determined, the OLS and GMM (and thereby GMM-GLS-IV) estimates are not consistent in general, but GLS-IV will remain consistent. For GMM and the first-step for GMM-GLS-IV we consider the set of instrumental variables Z={wtj,wthj,wthlj,j ≠ i,l=1,,h}. We use weekly data for 1-month forward rates and 3-month forward rates. In the former case, h=4 whereas in the second h=12. We consider the same data set as in subsection 5.1. We first present simulations tailored to this problem to shed light on the properties of the various estimators under a range of plausible scenarios.

We present two sets of Monte Carlo experiments. The DGP is inspired by the Hansen and Hodrick (1980) regression. We consider one-month forward rates so that h=4:

(20)

where yt+4=st+4ift,4i, i=UK and wtj=stjft4,4j, j={CAD,JP}. In the first set of Monte Carlo experiments, the simulation design is based on observed weekly spot and forward exchange rates for US-UK, US-CAD and US-JP. We use actual forecast errors wtj for US-CAD and US-JP. By construction, the regressors wtj will be exogenous. In subsection 6.1.2, we present the second set of Monte Carlo experiments, where the regressors wtj are jointly simulated with the error process ut, so they are serially correlated and non-exogenous.

We start with the case of exogenous instruments. We consider the DGP (20) under the null hypothesis H0:α0=0,β=0. We set α1=α2=1, so that wtj, j={CAD,JP} can be used as instruments. We use actual observed data for wtj, j={CAD,JP} and we generate yt+4 according with the DGP (20) imposing the null hypothesis. As discussed in Section 2, the “RE case” implies that ut+4 has an MA(3) representation. We simulate an MA(3) process with parameters calibrated to replicate the observed autocorrelation function up to lag 3 of the GLS-IV residuals from equation (19) using the real data set, following the same procedure as described in  Appendix A.3. The resulting parameters are as follows:

(21)

We simulate an error process based on equation (21), ut+h=θ(L)εt, where εti.i.dN(0,σε2) and σε2 is estimated using the GLS-IV residuals from the initial regression (20). We use the weekly spot exchange rates for US-UK, US-CAD, US-JP for the period between period November 2010 to April 2020 (492 observations). We perform 5,000 replications. We consider the following estimates: a) OLS applied to regression (20). We use HAC standard errors with the Quadratic Spectral weighting scheme of Andrews (1991) with automatic bandwidth selection using an ARMA(1,1) approximation; b) GMM: the estimate γ˜ (10) from regression (20) using the optimal weighting matrix as described in subsection 4.1.1. We use the set of instruments zt={wtj,wthj,wthlj,j={CAD,JP},l=1,,h}; c) GLS-IV: the estimate from the procedure described in subsection 4.2. For the Durbin regression (9), we set kmax=30. The autoregressive coefficients ρj are estimated using the efficient method described in Remark 2. For the IV estimates, we use the set of quasi-differenced instruments zt={wtj,wthj,j={CAD,JP}}. The variance estimate is as follows:

where X is the T×4 matrix of quasi-differenced regressors (including a constant term) and σ^GLSIV2 is the sample variance of the GLS-IV residuals u^GLSIV=ytβ^GLSIVyt1α^GLSIVwt; d) GMM-GLS-IV: the estimate from the procedure described in subsection 4.1.2. Step 1 uses the GMM estimate γ˜ described above. For the autoregression (11), we set kmax=30. For the IV estimate, we use the set of quasi-differenced instruments zt={wtj,wthj,j={CAD,JP}}. The variance estimate is computed as:

where X is the T×4 matrix of quasi-differenced regressors (including a constant term) and σ^GMMGLSIV2 is the variance of the GMM-GLS-IV residuals:

6.1.1 The case with exogenous instruments.

We start with the case with exogenous instruments and first assess the finite sample size of the estimators. The simulation results are presented in the first panel of Table 5, which report the MSE, bias, variance, coverage rate and average length of confidence intervals for the parameter β. Under the “RE case”, the error term follows an MA(3) process so that the OLS estimates are consistent. The GMM estimate has a variance that is half of the OLS variance but with a coverage rate below the nominal level. The GMM-GLS-IV procedure using the GMM as a first step estimate achieves an important reduction in MSE while maintaining coverage rates near the nominal level. The finite sample performance of GLS-IV and GMM-GLS-IV are similar. Both have the smallest MSE and yield confidence intervals with coverage rates near the nominal level and the shortest length, unlike GMM.

Table 5.

Simulation results with exogenous instruments. RE implies MA(3) errors; general implies ARMA(1,3) errors

CaseEstimator MSEBiasVarianceCoverageLength
REOLS0.193.430.170.930.16
GMM0.213.660.080.770.11
GMM-GLS-IV0.142.960.130.940.14
GLS-IV0.152.990.120.930.13
GeneralOLS3.6618.050.370.190.24
GMM0.796.980.310.750.22
GMM-GLS-IV0.092.390.110.960.12
GLS-IV0.102.420.110.950.12
Note(s):

Weekly data for US-CAD, US-JP for the period November 2010 to April; 2020 (T=492). For OLS we use HAC standard errors as described in the text

We now consider simulations under the “general case”. We use the DGP (20). However, we consider a departure of the efficient market hypothesis in which the error term ut+4 is serially correlated at lags i4. In particular, we assume that ut+4 follows an ARMA(1,3) process with MA coefficients given by (21) and AR coefficient ρ=0.6. The forward rate is generated in the same way as before. We consider the same family of estimators. The results are presented in the second panel of Table 5. In this case, as the error process is correlated beyond lag 3, OLS is no longer consistent. This translates in an important increase in MSE and bias. The confidence intervals are meaningless, in that they have huge size distortions, with a coverage rate smaller than 20%. Since the instruments are exogenous, the GMM estimate remains consistent in this case but is not efficient. The MSE of GMM is eight times that of FGLS. The GMM-GLS-IV estimate provides marked improvements over GMM and achieves important reduction in MSE along with confidence intervals having coverage rates near the nominal level. The finite sample performance of the GLS-IV estimate is similar to GMM-GLS-IV. Both estimates have the smallest MSE and variance with confidence intervals coverage rates near the nominal level and the shortest length.

We next consider the power of the various tests. We consider the DGP (20) for a grid values of β around the null hypothesis H0:β=0. In particular, we consider β{0.3,0.2,0.1,0,0.1,0.2,0.3}. We set α=(0,1,1), so that wtj, j={CAD,JP} can be used as instruments. We use the actual observed data for wtj, j={CAD,JP} and we generate yt+4 according with the DGP (20). The error process ut+4 is simulated as before: under the “RE case”, ut+4 follows an MA(3) process, and under the “general case” ut+4 follows an ARMA(1,3) process with the same parameter configurations used earlier. In Figures 3 and 4, we plot the empirical rejection frequencies of nominal α=0.05t-test of H0:β=0, for the same set of estimates considered in the previous section. Figure 3 pertains to the “RE case”, while Figure 4 present results for the “general case”. Note that for the “RE case” with MA(3) errors case, OLS has good power but is slightly outperformed by the FGLS-based estimates. Remarkably, GMM presents power distortions as it rejects the null hypothesis over 20%. In the “general case” with ARMA(1,3) errors, Figure 4 shows that the power of GMM-GLS-IV and GLS-IV remains the same while that of OLS exhibits huge power distortions skewed at negative values of β. The GMM estimate also has large size distortions and skewed power functions.

Figure 3.
A multi-line chart compares rejection frequency versus beta for G M M-G L S-I V, G L S-I V, O L S, and G M M methods.The multi-line chart displays rejection frequency on the y-axis against beta on the x-axis for 4 estimation methods, G M M-G L S-I V, G L S-I V, O L S, and G M M. All methods show high rejection frequencies close to 1 at beta values away from 0, while rejection frequency drops sharply near beta 0. The G M M-G L S-I V and G L S-I V curves overlap closely, reaching their minimum near 0.08 around beta 0. The O L S curve follows a similar pattern with slightly lower values near the minimum, while the G M M curve maintains comparatively higher rejection frequency around beta 0. The chart highlights strong symmetry around beta 0 and differing sensitivities among the estimation methods.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=0. Hansen and Hodrick regression, RE case with MA(3) errors

Figure 3.
A multi-line chart compares rejection frequency versus beta for G M M-G L S-I V, G L S-I V, O L S, and G M M methods.The multi-line chart displays rejection frequency on the y-axis against beta on the x-axis for 4 estimation methods, G M M-G L S-I V, G L S-I V, O L S, and G M M. All methods show high rejection frequencies close to 1 at beta values away from 0, while rejection frequency drops sharply near beta 0. The G M M-G L S-I V and G L S-I V curves overlap closely, reaching their minimum near 0.08 around beta 0. The O L S curve follows a similar pattern with slightly lower values near the minimum, while the G M M curve maintains comparatively higher rejection frequency around beta 0. The chart highlights strong symmetry around beta 0 and differing sensitivities among the estimation methods.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=0. Hansen and Hodrick regression, RE case with MA(3) errors

Close modal
Figure 4.
A multi-line chart compares rejection frequency versus beta for 4 statistical estimation methods.The line chart presents rejection frequency on the y-axis against beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M estimation methods. The G M M-G L S-I V and G L S-I V curves remain nearly identical, falling sharply to a minimum near beta 0 before rising again towards higher rejection frequencies at positive and negative beta values. The O L S curve differs substantially, increasing steadily from low rejection frequencies at negative beta values to nearly 1 at positive beta values. The G M M curve exhibits a smoother transition with moderate rejection frequencies around beta 0 and gradual increases towards the extremes.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=0. Hansen and Hodrick regression, general case with ARMA(1,3) errors

Figure 4.
A multi-line chart compares rejection frequency versus beta for 4 statistical estimation methods.The line chart presents rejection frequency on the y-axis against beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M estimation methods. The G M M-G L S-I V and G L S-I V curves remain nearly identical, falling sharply to a minimum near beta 0 before rising again towards higher rejection frequencies at positive and negative beta values. The O L S curve differs substantially, increasing steadily from low rejection frequencies at negative beta values to nearly 1 at positive beta values. The G M M curve exhibits a smoother transition with moderate rejection frequencies around beta 0 and gradual increases towards the extremes.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=0. Hansen and Hodrick regression, general case with ARMA(1,3) errors

Close modal

In summary, for the case with exogenous instruments, GMM-GLS-IV and GLS-IV clearly have better properties. The performance of OLS is nearly as good in the “RE case” but completely breaks down in the “general case”. Hence, GMM-GLS-IV and GLS-IV are clearly the more robust method of estimation and testing.

6.1.2 Simulations with non-exogenous instruments.

We now consider the case with non-exogenous instruments and first assess the finite sample size of the estimators followed by some power comparisons. The DGP is based on regression of Hansen and Hodrick (1980), considering one-month forward rates (h=4):

(22)

where yt+4=st+4ift,4i and wtj=stjft4,4j, j ≠ i. In this case, the set of instruments is simulated so that they are serially correlated and non-exogenous. Accordingly, we set:

for j=1,2, with vtji.i.d.N(0,1) independent of εti.i.d.N(0,1). Note that εt is shock affecting ut and thus, the instrument wtj is not exogenous whenever γ ≠ 0. We set γ=0.3 and ρw=0.5. The process yt+4 is generated according to the DGP (22) under the null hypothesis α0=0,β=0. We set α1=α2=1 so that wtj (j=1,2) can be used as instruments. Under the “RE case”, the error ut+4 follows an MA(3) process. Under the “general case”, ut+4 follows the same ARMA(1,3) process as in Section 6.1. We consider the same set of estimates: OLS + HAC, GMM, GLS-IV and GMM-GLS-IV. The sample size is T=300 and 5,000 replications are used.

The simulation results for the “RE case” are presented in the first panel of Table 6. In this case, the OLS estimates are consistent requiring only pre-determined regressors. It is, however, the less efficient estimate and the coverage rates of the associated confidence intervals are below the nominal level. The bad performance of GMM is again corrected when using the GMM-GLS-IV estimate; it achieves the smallest MSE and the coverage rate is better than that of the GLS-IV estimate. The GMM-GLS-IV and GLS-IV estimates have confidence intervals with coverage rates slightly below to the nominal level (92% and 87%, respectively). This results are in line with the Monte Carlo simulation results for the non-exogenous regressors and/or instruments cases in Perron and González-Coya (2022) and Perron and Olivari (2023).

Table 6.

Simulation results with non-exogenous instruments. RE implies MA(3) errors; general implies ARMA(1,3) errors

CaseEstimator MSEBiasVarianceCoverageLength
REOLS0.324.630.250.890.19
GMM0.374.840.190.820.17
GMM-GLS-IV0.173.280.130.920.14
GLS-IV0.233.740.130.870.14
GeneralOLS3.9618.460.470.260.26
GMM1.6710.630.560.650.26
GMM-GLS-IV0.193.470.150.910.15
GLS-IV0.263.710.150.890.14
Note(s):

Weekly data for US-CAD, US-JP for the period November 2010 to April; 2020 (T=492). For OLS we use HAC standard errors as described in the text

The simulation results for the “general case” with ARMA(1,3) errors are presented in the second panel of Table 6. Since the serial correlation in the errors extends beyond lag 3, OLS and GMM are no longer consistent. This is reflected in large MSE, bias and variance. The size distortions are exacerbated with coverage rates of the confidence intervals below 65% (26% for OLS). The finite sample performance of the GLS-IV estimate is in line with the fact that it is consistent. Surprisingly, despite the fact that the first step GMM estimates of the GMM-GLS-IV procedure are not consistent, the resulting GMM-GLS-IV estimate has smaller MSE than the GLS-IV estimate. These results suggest that the FGLS-based procedures are very robust to the first-stage estimates of the autocorrelation coefficients. The MSE of GLS-IV is almost 12 times smaller than the MSE of OLS, and it achieves with a variance that is on average five times smaller than the variance of the other estimates. The coverage rates of the confidence intervals of the GMM-GLS-IV and GLS-IV estimates are near the nominal level with the smallest length.

We now consider the power analysis. We consider the DGP (22) for a grid values of β around the null hypothesis H0:β=0. In particular, we consider β{0.3,0.2,0.1,0,0.1, 0.2,0.3}. We set α=(0,1,1), so that wtj, j=1,2 can be used as instruments. We generate yt+4 according with the DGP (22). Under the “RE case”, ut+4 follows an MA(3) process, and under the “general case” with ut+4 an ARMA(1,3) process. In Figures 5 and 6, we plot the empirical rejection frequencies of nominal α=0.05t-test of H0:β=0, for the same set of estimates considered in the previous section. Figure 5 pertains to the “RE case”, while Figure 6 pertains to the “general case with ARMA(1,3) errors.

Figure 5.
A multi-line chart compares rejection frequency versus beta for 4 estimation techniques.The chart illustrates rejection frequency on the y-axis versus beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M methods. The G M M-G L S-I V and G L S-I V curves overlap almost completely, decreasing from near 1 to their minimum around beta 0 before increasing symmetrically again. The O L S curve shows lower rejection frequencies near positive beta values compared with the other methods, while the G M M curve lies between the O L S and G L S-I V curves across most beta values. The overall pattern demonstrates a pronounced V-shaped behaviour centred around beta 0.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=0. Non-exogenous IVs, RE case with MA(3) errors

Figure 5.
A multi-line chart compares rejection frequency versus beta for 4 estimation techniques.The chart illustrates rejection frequency on the y-axis versus beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M methods. The G M M-G L S-I V and G L S-I V curves overlap almost completely, decreasing from near 1 to their minimum around beta 0 before increasing symmetrically again. The O L S curve shows lower rejection frequencies near positive beta values compared with the other methods, while the G M M curve lies between the O L S and G L S-I V curves across most beta values. The overall pattern demonstrates a pronounced V-shaped behaviour centred around beta 0.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=0. Non-exogenous IVs, RE case with MA(3) errors

Close modal
Figure 6.
A multi-line chart compares rejection frequency versus beta for 4 econometric methods.The line chart shows rejection frequency on the y-axis plotted against beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M methods. The G M M-G L S-I V and G L S-I V curves are nearly identical, exhibiting a deep minimum near beta 0 and rejection frequencies approaching 1 at larger absolute beta values. The O L S curve rises steadily from low rejection frequencies at negative beta values to high rejection frequencies at positive beta values. The G M M curve demonstrates moderate rejection frequencies across the range, increasing gradually from negative to positive beta values. The comparison highlights differing rejection behaviours and asymmetry among the estimation approaches.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=0. Non-exogenous IVs, General case with ARMA(1,3) errors

Figure 6.
A multi-line chart compares rejection frequency versus beta for 4 econometric methods.The line chart shows rejection frequency on the y-axis plotted against beta on the x-axis for G M M-G L S-I V, G L S-I V, O L S, and G M M methods. The G M M-G L S-I V and G L S-I V curves are nearly identical, exhibiting a deep minimum near beta 0 and rejection frequencies approaching 1 at larger absolute beta values. The O L S curve rises steadily from low rejection frequencies at negative beta values to high rejection frequencies at positive beta values. The G M M curve demonstrates moderate rejection frequencies across the range, increasing gradually from negative to positive beta values. The comparison highlights differing rejection behaviours and asymmetry among the estimation approaches.

Empirical rejection frequencies of nominal 5% t-Test of H0:β=0. Non-exogenous IVs, General case with ARMA(1,3) errors

Close modal

From the results in Figure 5 for MA(3) errors, all the tests have similar power functions, though somewhat lower for OLS and GMM when β is positive. The FGLS-based procedures exhibit the same power function, with a rejection frequency of the null hypothesis that is slightly higher than the 5% nominal level. For GMM the null rejection frequency is higher than 15%. The results in Figure 6 pertaining to the ARMA(1,3) errors case show that the power functions of the FGLS-based procedures exhibit the same behavior as in the MA(3) errors case. On the other hand, OLS and GMM are now subject to important size distortions. The power function of those estimates is biased toward negative values of β. Note that OLS rejects the null hypothesis H0:β=0 almost 10% of the times when β=0.2, while it rejects H0 with a frequency higher than 75% when β=0. Hence, the only reliable test are those obtained using the FGLS-based procedures.

Remark 3. Three consistent messages emerge across all Monte Carlo designs reported in Tables 1–6, spanning both the Fama and Hansen–Hodrick specifications, with and without exogenous instruments, and for one- and three-month forward rates. First, when the error process is MA(h1) – the rational expectations case – all estimators are consistent, but FGLS-based procedures achieve substantially smaller MSE and variance than OLS, with gains ranging from a factor of 3–5 depending on the specification. OLS with HAC standard errors delivers confidence intervals near the nominal coverage level in this case, but at the cost of considerably wider intervals and correspondingly lower power (Figures 1, 3 and 5). Second, when the error process exhibits serial correlation beyond lag h1 – the general case with ARMA(1,h1) errors – OLS becomes inconsistent, its MSE increases by an order of magnitude and its confidence intervals suffer severe size distortions with coverage rates falling below 20% in some designs. The FGLS-based procedures remain consistent and efficient, with coverage rates near the nominal level and MSE largely unchanged relative to the rational expectations case. Third, among the IV-based procedures required for specifications with lagged dependent variables, GMM-GLS-IV and GLS-IV deliver similar finite-sample performance, with both dominating GMM in terms of MSE, coverage and power. The overall practical implication is clear: the FGLS-based procedures are more efficient and at worst marginally more complex than OLS when the rational expectations hypothesis holds, and decisively superior when it does not, making them the dominant choice for h -step-ahead forecasting regressions in applied work.

We now report estimates of the model (19) using weekly spot exchange rates for the UK pound (US-UK), Canadian dollar (US-CAD) and Japanese yen (US-JP). We use the complete US-JP sample that spans from June 1995 to January 2023, with 1,442 observations. We consider the OLS + HAC, GMM, GMM-GLS-IV and GLS-IV estimates. For the FGLS-based methods we set kmax=40. The estimation results for one-month forward rates (h=4) are presented in Table 7, while those for three-month forward rates (h=12) are in Table 8. As documented in Section 7, the outcome of the Cumby and Huizinga (1992) test (CH-test) for autocorrelation at lags q>h1 show a rejection in all cases for the residuals of the Hansen and Hodrick (1980) regression (19), which opens the possibility of correlation between the regressors/instruments and the past errors.

Table 7.

Hansen and Hodrick (1980) one-month forward model estimation results (SE in parentheses)

 α0β
FXOLSGMMGMM-GLS-IVGLS-IVOLSGMMGMM-GLS-IVGLS-IV
US-UK−0.000 (0.001)0.001 (0.002)0.000 (0.001)0.000 (0.001)−0.087 (0.040)0.044 (3.875)1.436 (1.439)0.392 (0.278)
US-CAD−0.000 (0.001)0.000 (0.001)−0.000 (0.001)−0.000 (0.001)−0.035 (0.051)0.247 (1.012)0.040 (0.631)1.223 (0.959)
US-JP0.003 (0.002)0.004 (0.022)0.001 (0.001)0.007 (0.005)0.120 (0.054)0.024 (6.181)−0.039 (0.598)−0.534 (1.060)
α1α2
US-UK0.143 (0.059)0.029 (1.650)−0.442 (0.435)−0.081 (0.135)0.016 (0.045)0.025 (0.599)−0.112 (0.133)−0.037 (0.039)
US-CAD0.012 (0.046)−0.096 (0.280)0.053 (0.162)−0.455 (0.358)0.031 (0.037)0.050 (0.074)0.023 (0.024)0.034 (0.031)
US-JP−0.056 (0.063)−0.042 (1.262)−0.045 (0.092)0.050 (0.187)0.015 (0.058)0.013 (0.421)−0.027 (0.044)0.028 (0.045)
Note(s):

For US-UK, α1 is the coefficient for US-CAD, and α2 is the coefficient for US-JP; for US-CAD, α1 is the coefficient for US-UK and α2 is the coefficient for US-JP; for US-JP, α1 is the coefficient for US-UK and α2 is the coefficient for US-CAD

Table 8.

Hansen and Hodrick (1980) three-month forward model estimation results (SE in parentheses)

α0β
FXOLSGMMGMM-GLS-IVGLS-IVOLSGMMGMM-GLS-IVGLS-IV
US-UK−0.001 (0.004)0.000 (0.004)−0.000 (0.000)−0.001 (0.001)0.091 (0.077)0.937 (2.486)0.536 (0.424)0.376 (0.212)
US-CAD0.000 (0.003)0.000 (0.004)−0.000 (0.001)−0.000 (0.001)0.005 (0.071)0.103 (1.379)1.063 (0.604)0.594 (0.206)
US-JP−0.001 (0.004)−0.002 (0.006)0.000 (0.001)0.002 (0.004)0.104 (0.062)0.112 (0.703)1.090 (0.919)2.349 (1.372)
α1α2
US-UK0.055 (0.090)−0.378 (1.028)−0.102 (0.134)−0.104 (0.123)−0.026 (0.059)−0.146 (0.408)−0.067 (0.043)−0.054 (0.027)
US-CAD0.026 (0.070)−0.044 (0.420)−0.290 (0.158)−0.239 (0.097)0.097 (0.046)0.088 (0.130)0.017 (0.037)0.085 (0.023)
US-JP−0.051 (0.106)−0.038 (0.254)−0.193 (0.124)−0.342 (0.199)0.064 (0.108)0.111 (0.219)−0.074 (0.059)−0.062 (0.128)
Note(s):

For US-UK, α1 is the coefficient for US-CAD, and α2 is the coefficient for US-JP; for US-CAD, α1 is the coefficient for US-UK and α2 is the coefficient for US-JP; for US-JP, α1 is the coefficient for US-UK and α2 is the coefficient for US-CAD

In subsection 6.1.2, we provided evidence that the GMM-GLS-IV and GLS-IV estimates are consistent requiring only pre-determined instruments wt. We shall thus focus on the FGLS-based estimates. First note that both estimates are similar across all currencies and forecast horizons. For one- and three-month forward rates, the FGLS-based estimates cannot reject the null hypothesis H0:α0=β=α1=α2=0 for any of the currencies. In contrast, OLS rejects H0 for US-JP, for one- and three-month forward rates. There is no consensus in the literature about the rejection of this null hypothesis. For three-month forward rates (h=12), Hansen and Hodrick (1980) using OLS rejects the null hypothesis for US-CAD and two other currencies (Deutsche mark and Swiss franc) for data between 1975 and 1979.

Note the relatively high standard errors of the GMM estimates for all the currencies. This suggests that the instruments wtj may be non-exogenous. On the other hand, despite the fact that the first stage GMM estimate in the GMM-GLS-IV procedure is very noisy, the resulting GMM-GLS-IV estimate is much more efficient and close to the GLS-IV estimate. It can be argued that while the GLS-IV estimate is valid with only pre-determined instruments, it might still be subject to a “weak instrument” problem; see, e.g. Andrews et al., 2019. We provide statistical evidence to argue that this is not the case. In particular, we consider the test of Staiger and Stock (1997) for weak instruments and the Wu–Hausman exogeneity test (Hausman, 1978 and Wu, 1973). The detailed implementation for the GLS-IV estimate are outlined in  Appendix A.4.1 and  A.4.2. Note that standard weak instruments tests are valid for GLS-IV as the 2SLS regression (12) has uncorrelated errors. Table 9 reports the weak instruments test, the FSS test statistic and the exogeneity test FWH, see (A.5), together with the corresponding p-values. The null hypothesis for weak instruments is rejected for all currencies and h=4,12 at the 1% level of significance. The Wu-Hausman rejects the null hypothesis of exogeneity for US-UK and US-JP at least at the 5% level of significance for h=4,12 in all cases. Overall, these results indicate that we can be confident about the estimates and tests obtained using the GLS-IV procedure when applied to the Hansen and Hodrick (1980) regression (19).

Table 9.

Weak instruments (first row) and Wu–Hausman (second row) tests for GLS-IV

Forward ratesFX  df1 df2Statisticp-value
1-monthUS-UK72130913.022E-16
113797.910.00499
US-CAD7812997.7852E-16
113754.0220.044
US-JP48134917.902E-16
1139530.184.67E-08
3-monthUS-UK75128822.752E-16
1136156.271.13E-13
US-CAD75128824.6892E-16
1136110.670.001
US-JP39134818.7452E-16
113859.0660.00265

As shown in subsection 2.1, if the error term in regression (1) is serially correlated at lags q>k1, the OLS estimator is no longer consistent. In this section we briefly present the results from applying the Cumby and Huizinga (1992) test (CH-test) for autocorrelation at lags q>h1 to both the Fama (1984) regression (13) and the Hansen and Hodrick (1980) regression (19). This test is well suited for our purpose since the null hypothesis is that the error process is a moving average of known order q=h1>0 against the general alternative that the autocorrelations are nonzero at lags greater than q. The CH-Test is a Wald test of the null hypothesis that the regression error is uncorrelated with itself at lags q+1 through q+s. A general formulation for two-stage least squares and two-step two-stage least squares is presented in Cumby and Huizinga (1992). We require the errors ut to be unconditionally homoskedastic. We refer to the paper by Cumby and Huizinga (1992) and  Appendix A.4.3 for the details about the implementation of the test, which require the estimate of several quantities. We simply note that for the Fama regression (13) we use FGLS residuals:

where yt+h=st+hst and xt=ft,hst, whereas for the Hansen and Hodrick (1980) regression (19) we use GLS-IV residuals:

where yt+h=st+hifti and wtj=stjfthj for j ≠ i. In Table 10, we provide the results of the CH-Test statistics, labelled lq,s, of the null hypothesis that the regression error in the Fama regression (13) with 3-month forward rates is uncorrelated with itself at lags q+1 to q+s with q=h1=11 and s={12,15,20}. For both sample periods, we obtain a rejection of the null hypothesis for every s and conclude that the error term in regression (13) is serially correlated at lags q>11. Table 11 provide similar results for the Hansen and Hodrick (1980) regression (19). The specifications are similar except that we set s={5,10,50}. Again, for both sample periods, we have statistical evidence to reject the null hypothesis for every s. Hence, again here the error term in regression (19) is also serially correlated at lags q>11.

Table 10.

Cumby and Huizinga (1992) test of the null hypothesis that the Fama regression error is uncorrelated with itself at lags q+1 to q+s, q=11

PeriodLags US-UKUS-CAUS-JP
10/1984–01/202312196.499321.282223.436
1570.89448.85061.324
2057.01647.52644.952
10/1989–04/202112152.768280.233642.996
1534.91948.31529.323
2029.68075.37354.568
Table 11.

Cumby and Huizinga (1992) test of the null hypothesis that the Hansen–Hodrick regression error is uncorrelated with itself at lags q+1 to q+s, q=11

Period LagsUS-UKUS-CAUS-JP
10/1984–01/202351087.218203.5084646.771
101134.685219.5272624.1312
50525.2874330.38521115.21
10/1989–04/2021548811.7543522.5561825.54
10114112.9424805.465447.43
5093889.12377855.837567.25

Our FGLS framework applies to any setting in which a researcher estimates the parameters of a h-step-ahead linear forecasting equation using data sampled at a finer frequency than the forecast horizon. This structure arises naturally in a wide range of economic and financial applications beyond the UIP tests studied here. Examples include: (i) forecasting inflation or output growth at quarterly horizons using monthly data, where overlapping forecast errors induce MA(h1) serial correlation under rational expectations; (ii) predictive regressions for equity returns at horizons exceeding the sampling frequency, as in the stock return predictability literature; e.g. Campbell and Shiller (1988) and Stambaugh (1999); (iii) survey-based forecast evaluation, where professional forecasters issue multi-period-ahead predictions at regular intervals and the econometrician seeks to test forecast rationality or efficiency; and (iv) macro-finance term structure models in which bond risk premia are regressed on predictive variables at horizons longer than the observation frequency. In each of these settings, the key question is whether the forecast error is serially correlated only up to lag h1 – as implied by rational expectations – or beyond. If the latter, OLS is inconsistent when the regressors are not strictly exogenous, and the FGLS procedures developed here provide a consistent and efficient alternative.

We offer the following decision framework to guide practitioners in choosing among the estimators considered in this paper.

Step 1: Determine whether the regressors include lagged dependent variables. If the regression includes only contemporaneous or exogenous regressors, as in the Fama specification (13), the basic FGLS procedure of Section 3 applies directly. The Durbin regression is estimated via OLS, the autoregressive coefficients are used to quasi-difference the data, and the FGLS estimate is obtained from the quasi-differenced regression (7). No instrumental variables are required. This is the simplest and most efficient procedure when it applies.

Step 2: If lagged dependent variables are present, assess instrument exogeneity. When the regression includes lagged dependent variables, as in the Hansen-Hodrick specification (18), the choice between GMM-GLS-IV and GLS-IV depends on the plausible exogeneity of the available instruments. If the practitioner has strong reasons to believe the instruments are exogenous (i.e. uncorrelated with the error process at all leads and lags), GMM-GLS-IV is appropriate. It uses a first-stage GMM estimate to obtain consistent residuals, from which the autoregressive filtering coefficients are estimated. The advantage is that the first-stage GMM estimate and the associated residuals are consistent under exogeneity, yielding a clean identification of the autoregressive parameters.

However, if exogeneity of the instruments cannot be credibly maintained – and in many economic applications it cannot – GLS-IV should be preferred. GLS-IV requires only that the instruments to be pre-determined (uncorrelated with contemporaneous and future errors, but potentially correlated with past errors), a strictly weaker condition. Pre-determination is the natural assumption in most time series settings where instruments are lagged values of observable variables. As shown in our simulation experiments (Section 6.1.2), GLS-IV remains consistent and achieves competitive efficiency even with non-exogenous instruments, whereas GMM and GMM-GLS-IV can exhibit substantial bias and size distortions in this case. A notable finding from our simulations is that GMM-GLS-IV can perform well even when its first-stage GMM estimate is inconsistent (Table 6), suggesting that the FGLS filtering step is robust to moderate contamination of the first-stage residuals. Nevertheless, in the absence of a formal test confirming instrument exogeneity, we recommend GLS-IV as the default choice for its broader robustness guarantees.

Step 3: Diagnostic testing. Regardless of which estimator is selected, we recommend two diagnostic checks. First, the Cumby–Huizinga test (Section 7) should be applied to the estimated residuals to assess whether serial correlation extends beyond lag h1. If the null hypothesis of MA(h1) errors is not rejected, OLS with HAC standard errors remains consistent and the gains from FGLS are purely in efficiency. If the null is rejected, as we find in all our empirical specifications, the FGLS-based procedures are necessary for consistency. Second, when using GLS-IV, the Staiger–Stock weak instruments test and the Wu–Hausman exogeneity test (Section A.4) should be reported to assess instrument strength and to provide evidence on whether the IV correction is empirically relevant.

Step 4: Lag order selection. The maximum lag order p¯T for the BIC selection of pT should satisfy the rate condition p¯T3/T0 as T. In practice, we recommend p¯T=T1/3 as a starting point. For our sample sizes of approximately 1,500–2,000 weekly observations, this yields p¯T between 11 and 13, though we use the more generous values of p¯T=30 or 40 to allow the BIC sufficient flexibility. Practitioners working with shorter samples (e.g. monthly data with T<500) should use smaller values of p¯T to avoid overfitting. Following Ng and Perron (2005), the BIC comparison across different values of p should use the same effective number of observations to ensure a proper comparison.

We re-examined the statistical evidence about the hypothesis of UIP, which is a joint hypothesis of efficiency in the forward foreign exchange markets and rational expectations. Testing rationality hypothesis and exchange market efficiency is embedded in the general problem of estimating the parameters of a h-step-ahead linear forecasting equation. Under the null hypothesis, the forecast errors are serially correlated up to lags h1 and OLS is consistent. However, if the errors are serially correlated beyond lags h1, OLS is no longer consistent. This observation motivates using FGLS-based methods that are robust to the structure of the error process. We apply the FGLS procedure developed in Perron and González-Coya (2022) and we extend it to a setting with lagged dependent variables included as regressors. The resulting instrumental variables-based procedure, GLS-IV, is consistent requiring only pre-determined IVs. Using these FGLS methods, we study the main two UIP specifications in the literature: the Fama (1984) and Hansen and Hodrick (1980) regressions. We provide novel insights about the forward premium anomaly. Applying the consistent FGLS method we show that the estimates of β for three currencies are always non negative at the 1% significance level. A result that is contrary to the general finding that the OLS estimates are negative. Hence, the so-called “forward discount anomaly” is not as severe as previously thought. We also show statistically significant discrepancies between the OLS and GLS-IV estimates in the Hansen and Hodrick (1980) regression. We rationalize these discrepancies by showing that the regression residuals are in fact serially correlated beyond lags k, and thus OLS is not consistent, while the FGLS methods remain consistent. This point to the usefulness of adopting our more robust FGLS procedure. Not only is consistent and efficient under a wider range of contexts but, as we have shown, can deliver estimates that are different and point to a different assessment of the empirical facts.

The traditional interpretation of the forward discount anomaly rests on OLS estimates of β in the Fama regression that are negative, with an average around 0.88 across 75 published estimates Froot (1990). As Fama (1984) showed, this implies that the variance of the risk premium must exceed the variance of expected depreciation, a condition that has proven notoriously difficult to reconcile with standard asset pricing models. Much of the subsequent theoretical literature on habit formation, long-run risks and rare disasters has been motivated, at least in part, by the need to generate risk premia large enough to explain these extreme negative values. Our FGLS estimates reframe the problem. With β positive but generally below unity, the implied risk premium is still present but far more modest. The gap between β^FGLS and one is consistent with a small, positive covariance between the risk premium and the forward discount; this is in line in the sign and magnitude that calibrated consumption-based models can plausibly deliver without requiring extreme parameter values; see, e.g. Lustig and Verdelhan (2007) and Bansal and Shaliastovich (2013). In other words, the “puzzle” that much of the risk premium literature has sought to resolve was largely an artifact of inconsistent OLS estimation inflating the apparent size of the premium. The existing theoretical toolkit may already be adequate to explain the risk premia implied by our consistent estimates, rendering the anomaly considerably less anomalous than previously thought.

Andrews
,
D.W.
(
1991
), “
Heteroskedasticity and autocorrelation consistent covariance matrix estimation
”,
Econometrica
, Vol.
59
No.
3
, pp.
817
-
858
.
Andrews
,
I.
,
Stock
,
J.H.
and
Sun
,
L.
(
2019
), “
Weak instruments in instrumental variables regression: theory and practice
”,
Annual Review of Economics
, Vol.
11
No.
1
, pp.
727
-
753
.
Backus
,
D.K.
,
Gregory
,
A.W.
and
Telmer
,
C.I.
(
1993
), “
Accounting for forward rates in markets for foreign currency
”,
The Journal of Finance
, Vol.
48
No.
5
, pp.
1887
-
1908
.
Baillie
,
R.T.
,
Diebold
,
F.X.
,
Kapetanios
,
G.
and
Kim
,
K.H.
(
2023
), “
A new test for market efficiency and uncovered interest parity
”,
Journal of International Money and Finance
, Vol.
130
, p.
102765
.
Bansal
,
R.
and
Shaliastovich
,
I.
(
2013
), “
A long-run risks explanation of predictability puzzles in bond and currency markets
”,
Review of Financial Studies
, Vol.
26
No.
1
, pp.
1
-
33
.
Barberis
,
N.
,
Shleifer
,
A.
and
Vishny
,
R.
(
1998
), “
A model of investor sentiment
”,
Journal of Financial Economics
, Vol.
49
No.
3
, pp.
307
-
343
.
Bekaert
,
G.
and
Hodrick
,
R.J.
(
1993
), “
On biases in the measurement of foreign exchange risk premiums
”,
Journal of International Money and Finance
, Vol.
12
No.
2
, pp.
115
-
138
.
Berk
,
K.N.
(
1974
), “
Consistent autoregressive spectral estimates
”,
The Annals of Statistics
, Vol.
2
No.
3
, pp.
489
-
502
.
Bilson
,
J.F.O.
(
1981
), “
The “speculative efficiency” hypothesis
”,
The Journal of Business
, Vol.
54
No.
3
, pp.
435
-
451
.
Campbell
,
J.Y.
and
Shiller
,
R.J.
(
1988
), “
Stock prices, earnings, and expected dividends
”,
The Journal of Finance
, Vol.
43
No.
3
, pp.
661
-
676
.
Cumby
,
R.E.
and
Huizinga
,
J.
(
1992
), “
Testing the autocorrelation structure of disturbances in ordinary least squares and instrumental variables regressions
”,
Econometrica
, Vol.
60
No.
1
, pp.
185
-
195
.
De Grauwe
,
P.
and
Grimaldi
,
M.
(
2006
),
The Exchange Rate in a Behavioral Finance Framework
,
Princeton University Press
.
Durbin
,
J.
(
1970
), “
Testing for serial correlation in least-squares regression when some of the regressors are lagged dependent variables
”,
Econometrica
, Vol.
38
No.
3
, pp.
410
-
421
.
Engel
,
C.
(
1996
), “
The forward discount anomaly and the risk premium: a survey of recent evidence
”,
Journal of Empirical Finance
, Vol.
3
No.
2
, pp.
123
-
192
.
Fama
,
E.F.
(
1984
), “
Forward and spot exchange rates
”,
Journal of Monetary Economics
, Vol.
14
No.
3
, pp.
319
-
338
.
Froot
,
K.A.
(
1990
),
Short Rates and Expected Asset Returns
,
National Bureau of Economic Research Cambridge
,
Mass
.
Griliches
,
Z.
(
1961
), “
A note on serial correlation bias in estimates of distributed lags
”,
Econometrica
, Vol.
29
No.
1
, pp.
65
-
73
.
Hai
,
W.
,
Mark
,
N.C.
and
Wu
,
Y.
(
1997
), “
Understanding spot and forward exchange rate regressions
”,
Journal of Applied Econometrics
, Vol.
12
No.
6
, pp.
715
-
734
.
Hansen
,
L.P.
(
1982
), “
Large sample properties of generalized method of moments estimators
”,
Econometrica
, Vol.
50
No.
4
, pp.
1029
-
1054
.
Hansen
,
L.P.
and
Hodrick
,
R.J.
(
1980
), “
Forward exchange rates as optimal predictors of future spot rates: an econometric analysis
”,
Journal of Political Economy
, Vol.
88
No.
5
, pp.
829
-
853
.
Hausman
,
J.A.
(
1978
), “
Specification tests in econometrics
”,
Econometrica
, Vol.
46
No.
6
, pp.
1251
-
1271
.
Liviatan
,
N.
(
1963
), “
Consistent estimation of distributed lags
”,
International Economic Review
, Vol.
4
No.
1
, pp.
44
-
52
.
Lustig
,
H.
and
Verdelhan
,
A.
(
2007
), “
The cross section of foreign currency risk premia and consumption growth risk
”,
American Economic Review
, Vol.
97
No.
1
, pp.
89
-
117
.
Malinvaud
,
E.
(
1966
),
Statistical Methods of Econometrics
,
North-Holland
,
Amsterdam
.
Mankiw
,
N.G.
and
Reis
,
R.
(
2002
), “
Sticky information versus sticky prices: a proposal to replace the new Keynesian Phillips curve
”,
The Quarterly Journal of Economics
, Vol.
117
No.
4
, pp.
1295
-
1328
.
Muth
,
J.F.
(
1960
), “
Optimal properties of exponentially weighted forecasts
”,
Journal of the American Statistical Association
, Vol.
55
No.
290
, pp.
299
-
306
.
Ng
,
S.
and
Perron
,
P.
(
2005
), “
A note on the selection of time series models
”,
Oxford Bulletin of Economics and Statistics
, Vol.
67
No.
1
, pp.
115
-
134
.
Perron
,
P.
and
González-Coya
,
E.
(
2022
), “Feasible GLS for time series regression”, ”
Working Paper
,
Department of Economics, Boston University
.
Perron
,
P.
and
Olivari
,
M.
(
2023
), “GLS-IV in time series regressions”, ”
Working Paper
,
Department of Economics, Boston University
.
Schwarz
,
G.
(
1978
), “
Estimating the dimension of a model
”,
The Annals of Statistics
, Vol.
6
No.
2
, pp.
461
-
464
.
Sims
,
C.A.
(
2003
), “
Implications of rational inattention
”,
Journal of Monetary Economics
, Vol.
50
No.
3
, pp.
665
-
690
.
Small
,
D.S.
(
2007
), “
Sensitivity analysis for instrumental variables regression with overidentifying restrictions
”,
Journal of the American Statistical Association
, Vol.
102
No.
479
, pp.
1049
-
1058
.
Staiger
,
D.
and
Stock
,
J.H.
(
1997
), “
Instrumental variables regression with weak instruments
”,
Econometrica
, Vol.
65
No.
3
, pp.
557
-
586
.
Stambaugh
,
R.F.
(
1999
), “
Predictive regressions
”,
Journal of Financial Economics
, Vol.
54
No.
3
, pp.
375
-
421
.
Wallis
,
K.F.
(
1967
), “
Lagged dependent variables and serially correlated errors: a reappraisal of three-pass least squares
”,
The Review of Economics and Statistics
, Vol.
49
No.
4
, pp.
555
-
567
.
Woodford
,
M.
(
2003
),
Interest and Prices: Foundations of a Theory of Monetary Policy
,
Princeton University Press
.
Wu
,
D.M.
(
1973
), “
Alternative tests of independence between stochastic regressors and disturbances
”,
Econometrica
, Vol.
41
No.
4
, pp.
733
-
750
.

Assume that yt has a permanent and a transitory component, yt=y¯t+ωt, where the permanent component is defined by y¯t=y¯t1+εt=i=1tεi, where εii.i.d(0,σε2), ωii.i.d(0,σω2) and εi,ωi are assumed to be independent. Then we can write:

Hence:

Thus, we can write the forecast error as ut=(1κ)ut1+ξt, where ξt=εt+ωtωt1. Note that εt and ωt can be allowed to be correlated, so that ξt is invertible in general.

A.2 Efficient estimate of autoregressive coefficients for GLS-IV

Suppose that n=2. The efficient linear combination ρ˜j=λρ˜1j+(1λ)ρ˜2j is obtained when λ minimizes Var(ρ˜j). Hence, the minimization problem is minλ1,λ2Var(λ1ρ˜1j+λ2ρ˜2j), subject to λ1+λ2=1.The first-order conditions of this problem are:

where μ is the Lagrangian multiplier. The solution is:

A.3 Simulation design

We simulate an MA(h1) process with parameters calibrated to replicate the observed autocorrelation function up to lag h1 of the FGLS residuals of regression (13) using US-UK data for the Fama regression (Section 5) and using GLS-IV residuals of regression (19) for the Hansen and Hodrick (1980) regression (Section 6). We obtain the parameters θ1,,θk1 by solving the non-linear system of h1 equations given by the MA(h1) autocorrelation functions (ACF) for j=1,,h1. Let vt be an MA(q) process with q=h1, then the variance of vt is as follows:

The autocovariance function of vt is as follows:

Thus, the autocorrelation function of vt is as follows:

Using the observed first q autocorrelations of the residuals of the regression, ACFj (j=1,,q) we define a system of q non-linear equations with q unknown variables θ^1,,θ^q:

A.4 Diagnostic tests for GLS-IV

Rewrite the model using the quasi-differenced variables yt, ytk and wt as:

(A.1)
(A.2)

where (A.1) is the structural equation of interest, y=(yk+1,,yT) is a (Tk)×1 vector and Y=[yk] is a (Tk)×1 vector with the lagged dependent variable (the only endogenous variable in the model). (A.2) is the reduced form equation for Y, W is the (Tk)×K1 matrix of exogenous regressors with row t, Wt=[1,wt1,wt2] and K1=3 (i.e. W includes a constant term). Z is the (Tk)×K2 matrix of quasi-differenced instruments with row t, Zt={wtj,wtkj,j ≠ i} and K2=4. ε and V are, respectively, a (Tk)×1 vector and a (Tk)×2 matrix of error terms. Note that the quasi-differenced regression (A.1) has serially uncorrelated errors, with covariance matrix Σ. Assume that E[εt2]=σεε,E[Vtεt]=ΣVε and E[VtVt]=ΣVV. Let Z¯=[X,Z], it is assumed throughout that E[Z¯t(utVt)]=0.

A.4.1 Staiger and Stock (1997) weak instruments test

We are interested in testing Π=0 in the regression (A.2). Π shall be modeled as local to zero, so that the F statistic is Op(1). Staiger and Stock (1997) make the assumption that Π=ΠT=c/Tk where c is a fixed K2×1 vector. Before proceeding we provide some additional definitions and notation. Let Q=E[Z¯tZ¯t], partitioned so that E[WtWt]=QWW, E[WtZt]=QWZ and E[ZtZt]=QZZ. Also let ρ=ΣVV1/2ΣVεσεε1/2. Let PR=R(RR)1R and MR=IPR where R is a general a×b matrix with ab, and let “” denote the residuals from the projection on W, so Z=MWZ, Y=MWY, etc. Let W¯=[Y,W] and Y¯=[y,Y] and let Ik denote the k-dimensional identity matrix. Staiger and Stock (1997) assume that the following limits hold jointly: 1) (εε/T,Vε/T,VV/T)p(σεε,ΣVε,ΣVV); 2) T1W¯W¯pQ; 3):

where ΨΨWε,ΨZε,vec(ΨWV), with ΨWVN(0,ΣQ). Define λ=Ω1/2CΣVV1/2, where Ω=QZZQZXQXX1QXZ:

and zV=Ω1/2(ΨZVQZXQXX1ΨXV)ΣVV1/2. The random variable [zuvec(zV)] is distributed N(0,Σ¯IK2), where Σ¯ is the (n+1)×(n+1) matrix with Σ¯11=1,Σ¯22=In,Σ¯12=ρ and Σ¯21=ρ, where Σ¯ is partitioned conformably with Σ. Finally, let:

(A.3)

and:

(A.4)

The 2SLS estimate of (β,γ) is:

By standard projection arguments, the 2SLS estimate of β is as follows:

The Wald statistic testing Π=0 is W=tr(GT)/K2, where GT=Σ^VV1/2YPZYΣ^VV1/2, with Σ^VV=YMZ¯Y/(TK1K2). Staiger and Stock (1997) show that the limit distribution of GT is ν1 defined in (A.3). As we just have one endogenous variable, ytk, the F statistic FSS=GT/K2 converges to a non-central χK22/k2 with noncentrality parameter λλ. In the general case, with more than one endogenous variable, λλ is the matrix of noncentrality parameters of the limiting noncentral Wishart random variable ν1.

A.4.2 Wu–Hausman test of exogeneity

The Wu–Hausman (WH) test (see Hausman, 1978 and Wu, 1973) examines the null hypothesis that Y is exogenous (i.e. p=0) by checking for a statistically significant difference between the OLS and 2SLS estimates of β. The test statistic is as follows:

With:

Its limit distribution is as follows:

where Δ0(0)=ν11ν2 and S1(b)=12ρb+bb. Under the null hypothesis ρ=0, FWH simplifies to:

(A.5)

where ζ=ν11/2(λ+zV)ηN(0,In) with η=(zuzVρ)/1ρρ and ζ and ν1 are independent. Note that since ζζ/(1+ζν11ζ) ≤ζζχn2, applying χn2 critical values to FWH results in asymptotically conservative tests. However, as noted by Staiger and Stock (1997), a size adjustment of FWH is infeasible because the distribution depends on λλ/K2.

A.4.3 Testing for serial correlation at lags q > k−1

We describe in some details the test of Cumby and Huizinga (1992) for autocorrelation structure of the OLS residuals (CH-Test). This test is perfectly suited for our purpose as it allows to have under the null hypothesis a regression error process with a moving average of known order q=k1>0 against the general alternative that the autocorrelations of the regression error are nonzero at lags greater than q. The CH-Test is a Wald test of the null hypothesis that the regression error is uncorrelated with itself at lags q+1 through q+s. Consider the general formulation of the CH-Test for an OLS regression. Here we present the Cumby and Huizinga (1992) test for autocorrelation structure of the OLS residuals of equation (1) with no instrumental variables. A general formulation for two-stage least squares and two-step two-stage least squares is presented in Cumby and Huizinga (1992). The model is as follows:

where Xt is a vector of the n scalar predetermined regressors. The regression errors, ut are assumed to be serially correlated up to a known lag q0 and their autocorrelations at all lags greater than q are required to be zero under the null hypothesis. We require the errors ut to be unconditionally homoskedastic. The CH-Test statistic is as follows:

where V^r,B^,V^d,C^,D^ are consistent estimates of Vr,B,Vd,C,D. Here, r is a s×1 vector r=[rq+1,rq+2,,rq+s]:

Vd is the asymptotic covariance matrix of the estimator β^:

with Ω=limTT1E[uu]. Let D be the k×k matrix D=plimTT(XX)1. B is the s×k matrix with i, jth element:

Let ξi,t=ututqi for i=1,,s,ωj,t=utXj,t for j=1,,k, the ijth element of the s×s matrix Vr be given by Vr(i,j)=σu4n=qqE(ξi,tξj,tn), and the ijth element of the s×h matrix C be given by C(i,j)=σu2n=qqE(ξi,tωj,tn). To consistently estimate the test statistic lq,s, we need consistent estimate of the errors ut. For the Fama regression (13), we use FGLS residuals:

where yt+k=st+kst and xt=ft,kst. Whereas for the Hansen and Hodrick (1980) regression (19), we use GLS-IV residuals:

where yt+k=st+kifti and wtj=stjftkj for j ≠ i. We estimate Ω as discussed in subsection 4.1.1.

Published in Applied Economic Analysis. Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence maybe seen at Link to the terms of the CC BY 4.0 licenceLink to the terms of the CC BY 4.0 licence.

or Create an Account

Close Modal
Close Modal