The evidence

Not owning the wrong companies beats picking the right ones.

That is the result the method rests on, and it held in both halves of a 28.5-year record while the more obvious version of the idea — buy the best companies — quietly stopped working.

It came out of 151 tested strategies, 13 of which survived. Below is what did, the bar they had to clear, and then a register of ones that did not — the best known of them, not all 138.

151
strategies tested
13
cleared correction for multiple testing
13
of those 13 also held on the years held back
14
failures named in full, with what was run

Every rule was specified on 1998–2014 and measured once on 2015–2026. A result that only exists on the years it was built from is not a result.

What you avoid matters more than what you pick

The weakest companies on cash generation underperformed the rest by about four points a year — and by almost exactly the same margin in 1998–2014 as in 2015–2026. The strong end of the ranking did not hold up nearly as well: it worked in the first half and faded in the second, as the index came to be carried by a few very large companies these measures do not select.

Nothing here involves betting against those companies — most of them still rise, just by less. They are simply left out, which is something any private investor can do.

See the companies left out

Cash beats reported earnings

Cash produced by the business was the strongest of the 108 measures tested. Reported profit depends on judgement calls — when a sale is booked, how quickly an asset is written down — and two companies with identical cash can report very different profits. Cash collected leaves less room for that. The accounting adjustments that explain the gap between the two added nothing once cash itself was accounted for.

It is why the screen reads the cash flow statement, not the income statement.

How fast a company files predicts how it performs

The days between a quarter ending and the results being filed is behavioural, not accounting. A company slower than its own history is having trouble closing its books. It was the strongest single factor of everything tested, stronger in the years held back than the years fitted on, and uncorrelated with every other measure in the method.

It is the only one of the four that a competitor could not read straight off the accounts — though if enough people started trading on it, that would change. The portfolios do not depend on it alone: the three cash measures carry most of the weight, and they were strong in both halves of the record on their own.

Concentration is where a backtest lies to you

Take a very concentrated version of this screen — six companies — and run it six times over, changing nothing but which month it rebalances in. The six runs finish as much as 6.8 percentage points a year apart. Same rules, same companies available, and the month alone decides the answer. That is luck wearing the costume of a strategy. The split test says the same thing from the other side: an eight-company book beat the index by 12.4 points a year on the data it was designed against, then finished behind it on the years set aside.

Every size below twenty companies failed on the years set aside. That is why the tightest book here holds twenty-five and not eight.

How every claim here was tested

Point in time

A company's figures are used only once they were actually filed with the regulator, plus sixty days. Nothing in a test can know something the market did not.

Failures included

The universe holds companies that later went bust or were taken over. A holding that stops trading is booked at a 30% loss instead of quietly vanishing from the record.

Held back and looked at once

Every rule was designed on 1998–2014 and measured on 2015–2026 exactly once. A result that only exists on the years it was built from is not a result.

Corrected for searching

Testing many ideas guarantees some look good by chance. Where many were tried together, the bar for calling anything significant was raised in proportion.

The failure register

Publishing what worked is easy. This is the other half: the recognisable strategies that were tested here and did not survive, each with the specification that was actually run and what the outcome does not prove.

Much of what follows is other people's published research. A strategy that fails here failed on this universe, over this window — which is a statement about this test and not about the original work. Anything where our own specification departed materially from the published one is not listed at all: a failed implementation says nothing about the strategy, and hedging that in a footnote is worse than leaving it out.

Did not replicate · 10
Ran close to the published specification and the effect was not there on this universe and window.
Not measurable here · 2
No spread either way — the data could not separate the effect from noise.
True, and not worth trading · 2
The pattern is really there in the data, and acting on it still made less money than simply holding the index once costs were paid.
12-1 month price momentumDid not replicate
Source

Jegadeesh & Titman (1993), Journal of Finance

The claim

Stocks that rose over the past twelve months, skipping the most recent month, keep outperforming over the following month.

What we ran

The price one month ago divided by the price twelve months ago, minus one, so that the most recent month is deliberately skipped. Companies sorted into five equal groups on that measure, with the expected direction declared before the test was run, held one month, drawn from the 1,000 largest US companies by market value excluding financials and utilities, before trading costs. The reported figure is the return of the strongest fifth minus the weakest fifth.

Window

Monthly, 1998-2026 - 331 monthly observations. Reported separately for 1998-2014 and for 2015-2026, the later stretch having been set aside before testing began and not examined while factors were being chosen.

Result

The direction and the size of the published effect, without the statistical significance. The gap between the strongest and weakest fifths was 6.69% a year with the same amount in each company and 5.06% weighted by company size, positive in 56.8% of months, and it ran +6.23% over 1998-2014 and +7.35% over the held-back 2015-2026 - one of the few factors in the round that was STRONGER on data it had not been chosen on. But the measurement is only 1.39 times its own margin of error (t-statistic 1.39, p-value 0.165 against a conventional 0.05 threshold), so it cleared neither our significance screen nor the correction for having tested 34 factors at once, and it was not carried into the portfolio.

What this does not show

This is a failure of statistical power and of our own pre-set bar, not evidence that momentum is absent - the estimate is large, positive and consistent across both halves of the record, and we say so plainly. Momentum's month-to-month swings are wide enough that 28 years of monthly data on 1,000 large companies cannot resolve it at conventional significance. We also charged no trading costs here, which a real momentum portfolio would pay heavily.

The accruals anomalyDid not replicate
Source

Sloan (1996), The Accounting Review

The claim

Companies whose earnings consist largely of accounting entries, not cash actually collected, subsequently underperform.

What we ran

Net income minus cash flow from operations, divided by total assets - the cash-flow-statement form of the measure - with the expected direction declared before the test was run, so that companies with high accruals were expected to underperform. Companies sorted into five equal groups, held one month, drawn from the 1,000 largest US companies by market value excluding financials and utilities, before trading costs. The reported figure is signed so that a positive number means the published direction held.

Window

Monthly, 1998-2026 - 342 monthly observations. Reported separately for 1998-2014 and for 2015-2026, the later stretch having been set aside before testing began and not examined while factors were being chosen.

Result

In the predicted direction, but short of our bar. The gap between the extreme fifths ran 2.70% a year with the same amount in each company and 5.26% weighted by company size, and it was positive in both halves of the record: +2.96% over 1998-2014 and +2.32% over the held-back 2015-2026. But the measurement is 1.69 times its own margin of error (t-statistic 1.69, p-value 0.091 against a conventional 0.05 threshold), and it did not survive the correction for having tested 34 factors at once.

What this does not show

Our figure is consistent with the published direction, and with the literature's own account of the accrual effect weakening after 2000, so this is a near-miss, not a contradiction. It is also consistent with Ball, Gerakos, Linnainmaa & Nikolaev (2016), whose claim is that cash-based operating profitability absorbs the accrual effect - and cash-based operating profitability was the single strongest of the 90 factors we have tested. Sloan's original measure is built from the balance sheet, not the cash flow statement, so ours is the modern implementation and not his.

The asset-growth anomalyDid not replicate
Source

Cooper, Gulen & Schill (2008), Journal of Finance. Our own test ranks companies on plain asset growth and does not follow the paper's construction.

The claim

Companies that expand their balance sheets fastest earn the lowest subsequent returns.

What we ran

Total assets divided by total assets four filings earlier, minus one. Companies sorted into five equal groups on that growth rate, held one month, drawn from the 1,000 largest US companies by market value excluding financials and utilities, before trading costs, with the direction declared in advance so that the fastest expanders were expected to underperform. A three-year version of the same measure was tested separately in a later round.

Window

Monthly, 1998-2026 - 342 monthly observations. Reported separately for 1998-2014 and for 2015-2026, the later stretch having been set aside before testing began and not examined while factors were being chosen.

Result

The gap between the slowest- and fastest-expanding fifths was +3.25% a year when every company counts equally, but -0.77% once the groups are weighted by company size, and the measurement is only 1.14 times its own margin of error (t-statistic 1.14, p-value 0.257 against a conventional 0.05 threshold). It ran +4.33% over 1998-2014 and +1.65% over the held-back 2015-2026. So: directionally as published when every company counts equally, gone once weighted by company size, and short of the correction for having tested 34 factors at once.

What this does not show

The gap between the equal-amount figure (+3.25%) and the size-weighted one (-0.77%) is the whole story: on our universe the effect lives in the smaller companies, and the published result is reported as strongest in small stocks. Our test therefore excludes precisely where the effect is documented to be largest, and a null on the 1,000 largest US companies is not evidence against it.

The Best Six Months ("Sell in May and go away")Did not replicate
Source

Stock Trader's Almanac 2025 (Hirsch & Mistal); the Best Six Months switching strategy was published in 1986

The claim

Almost all of the market's gain arrives between November and April, so an investor can sit out May to October and give up nothing.

What we ran

A daily index of the largest 500 US companies, held only from November through April and switched to cash from May through October. The index is our own construction, not a licensed one: the 500 largest US companies by market value, excluding financials and utilities, weighted by company size, built from prices adjusted for dividends and splits, with the basket reset once a month and held unchanged within the month. Each switch into or out of the market is charged 10 basis points, that is 0.10% of the amount traded. Because the answer depends on what the idle cash earns, the whole test is repeated with cash paying 0%, 2% and 4% a year. Three tests were fixed in advance, and the rule had to clear all three. First, beat simply buying and holding the same index over the same window. Second, survive a correction for the fact that many rules were being tested at once, so that a rule cannot qualify merely by being the luckiest of the batch: this is the standard Benjamini-Hochberg correction, set to tolerate one false positive in ten. Third, be positive in both halves of the record, not only overall. One further check was also run, though it was not one of the three bars: whether holding the market on these particular days beat holding it for the same number of days chosen at random, using 4,000 random schedules with the days shuffled inside each calendar year. That check asks whether the calendar matters or only the amount of time spent invested.

Window

Daily, 2 January 1998 to 28 August 2026 - 7,208 trading days. The two halves used for the both-halves test split at 1 January 2013. No part of the record was held back untested in this round.

Result

Failed. Buying and holding the same index over the same window compounded at 7.84% a year, and the seasonal switch did not beat it. Of the 21 Almanac rules tested in this round - 18 that decide when to be invested and 3 that pick individual companies - none cleared all three bars, and every one of the timing rules came out negative over 2013-2026, which fails the both-halves bar on its own. One detail of the correction is worth stating precisely: it was applied across the 18 timing rules together, while the 3 company-picking rules were judged separately against their own resampling tests instead of being pooled in, so the correction was never computed over all 21 at once. None of those three cleared even its uncorrected bar against the fairer, equal-amount benchmark. The per-rule figures were displayed on screen as the test ran and never written to a results file, so what survives for this particular rule is the round-level verdict and not a saved number of its own.

What this does not show

This says nothing about the Almanac's own record, which begins decades before 1998. Our window sits entirely after the strategy was published, and it contains no extended bear market inside the May-October half - so any rule that spends that half in cash loses here for mechanical reasons, not because the seasonal pattern is absent.

Best Six Months with MACD timingTrue, and not worth trading
Source

Stock Trader's Almanac 2025 (Hirsch & Mistal) - the MACD-timed variant of the Best Six Months switching strategy

The claim

Waiting for MACD to turn bullish before entering the November-April window, and for it to turn bearish before leaving, improves on the fixed-date seasonal switch.

What we ran

The same daily index and the same cost and cash assumptions as the plain Best Six Months switch: our own index of the 500 largest US companies by market value, excluding financials and utilities, weighted by company size, 10 basis points (0.10% of the amount traded) per switch, and cash credited at 0%, 2% and 4% a year. The rule enters the market when the calendar month is October to March and the MACD trend gauge sits above its signal line, and leaves when the month is April to September and the gauge sits below it. MACD is a standard trend indicator: the gap between a 12-day and a 26-day average of price, judged against a 9-day average of that same gap. The reading is deliberately lagged by one day, because it is computed from the closing price and so cannot decide the position that earns the same day's return. Judged on the same three tests fixed in advance as every other rule in the round: beat buying and holding the same index, survive the correction for having tested 18 rules at once, and be positive in both halves of the record.

Window

Daily, 1998-2026 - 7,208 trading days. The two halves used for the both-halves bar split at 1 January 2013.

Result

The closest thing to a survivor in the whole seasonal programme, and still a failure on our bar. It beat buying and holding the index when cash pays 2% a year, by 0.59 percentage points a year, and when cash pays 4%, by 1.68 points. It was also markedly steadier: 0.62 units of annual return for each unit of annual price movement against 0.41 for buy-and-hold, where higher is calmer, and a worst peak-to-trough fall of -35.5% against -57%. And holding the market on its chosen days beat holding it for the same number of days chosen at random, with a p-value of 0.029 - p being the probability of a gap at least this large arising by chance if the calendar made no difference, and 0.05 the conventional threshold below which a result is called significant. But it did not survive the correction for having tested 18 rules at once, and it was negative over the 2013-2026 half, so it failed two of the three bars set in advance. It is a way of reducing risk, not a way of earning more.

What this does not show

Its advantage comes overwhelmingly from being out of the market in 2000-2002 and in 2008; the second half of our record contains no comparable fall, so this is as much a test of one market regime as of the rule. Two further notes concern our own test, not the Almanac's claim. The steadiness figure quoted above is annual return divided by annual volatility, not return in excess of cash divided by volatility, so it is not a Sharpe ratio and should not be read against published Sharpe ratios. And an earlier version of this test read the trend gauge on the same day it was computed, which showed 21-23% a year on the later period against 12% for buy-and-hold; that was a look-ahead error in our own code, which we found and corrected, and the figures above come from the corrected version.

The book-to-market value premiumNot measurable here
Source

Fama & French (1992), Journal of Finance. Our own test ranks companies on plain book value divided by price and does not follow the construction in any one paper.

The claim

Stocks that are cheap relative to book value outperform expensive ones.

What we ran

Common shareholders' equity divided by the company's market value on the day, with the expected direction declared before the test was run. Companies sorted into five equal groups on that ratio, held one month, drawn from the 1,000 largest US companies by market value excluding financials and utilities, before trading costs; the reported figure is the return of the cheapest fifth minus the dearest fifth. A second attempt applied the standard repair for the fact that book value no longer captures much of what a modern company is worth. Research and development spending was accumulated as an asset and written down at 30% a year, and 30% of selling, general and administrative expense was accumulated as organisational capital and written down at 20% a year. That accumulated stock was then added to book equity before dividing by market value.

Window

Monthly, 1998-2026 - 342 monthly observations. Reported separately for 1998-2014 and for 2015-2026, the later stretch having been set aside before testing began and not examined while factors were being chosen.

Result

Absent. The plain version's gap between the cheapest and dearest fifths was -0.85% a year with the same amount in each company and -3.99% weighted by company size, a measurement well under one times its own margin of error (t-statistic -0.26, p-value 0.798 against a conventional 0.05 threshold) - which is to say indistinguishable from zero, not reliably negative. Across the record it ran +0.67% then -3.10%. The intangibles repair did not rescue it either: t-statistic 0.35, p-value 0.724, +3.57% then -2.07%. Worth noting alongside that: how intangible-INTENSIVE a company is did work as a quality signal in the same round, at a t-statistic of 2.90. It is specifically the attempt to restore book value as a measure of cheapness that did not work here.

What this does not show

1998-2026 on large US companies excluding financials is close to the least favourable window and universe available for this premium: it removes the small companies and the financial sector where book value is most informative, and it spans a widely documented drought for value investing. A plain sort into five groups is also not the construction Fama and French use - their value factor is a size-balanced portfolio rebalanced once a year, long the cheap companies and short the expensive ones. The absence here is a fact about our slice of the data, not about their sample.

The Free Lunch (52-week lows bought in mid-December)Did not replicate
Source

Stock Trader's Almanac 2025 (Hirsch & Mistal)

The claim

Stocks sitting at their 52-week lows in mid-December have been driven there by tax-loss selling and rebound sharply into February.

What we ran

From the 2,500 largest US companies by market value, excluding financials and utilities, buy every company trading within 2% of its lowest price of the past 252 trading days - about one year. Positions are opened on the first trading day on or after 15 December with the same amount in each, at least five companies must qualify, and the basket is held to the first trading day on or after 15 February. Measured against two benchmarks over the identical window: the largest 500 companies weighted by company size, and the same 500 held in equal amounts. Significance came from 4,000 resamplings of the individual years.

Window

One observation a year: each mid-December to mid-February window from 1998/99 through 2025/26, about 28 in all.

Result

Ahead of the size-weighted 500 by 2.98% a year, with a p-value of 0.080 - short of the conventional 0.05 threshold, p being the probability of a gap at least this large arising by chance if there were no effect. Against the same 500 held in equal amounts it was ahead by only 1.84%, with a p-value of 0.252. That second comparison is the fair one for a basket of beaten-down companies held in equal amounts, and on it the result does not reach significance: most of the apparent edge is the equal-position-size and smaller-company tilt, not the beaten-down selection itself.

What this does not show

Ours is a mechanical within-2%-of-the-one-year-low screen run over the largest 2,500 US listings; the Almanac's published Free Lunch list is compiled differently, so the two baskets are not the same basket. With about 28 annual observations, the range of values consistent with our data comfortably contains both the Almanac's number and zero.

The holiday effectTrue, and not worth trading
Source

Stock Trader's Almanac 2025 (Hirsch & Mistal)

The claim

The trading days surrounding the market's holidays are unusually strong.

What we ran

Hold the daily index only on days within three trading days of a market holiday - holidays being identified as weekdays missing from the trading calendar - and hold cash otherwise, with 10 basis points (0.10% of the amount traded) charged per switch and cash credited at 0%, 2% and 4% a year. Because sitting in cash is itself costly, the same idea was also tested in a form that never leaves the market: double exposure to the index on holiday windows, normal exposure the rest of the time, with the borrowing charged at 3% a year. That version is scored against constant exposure at the same average level, so that only the timing, and not the extra exposure, can be credited.

Window

Daily, 1998-2026. For the switching version, the two halves used for the both-halves bar split at 1 January 2013. For the double-exposure version, the rule was selected on 1998-2014 and then measured on 2015-2026, a stretch set aside and not examined while the rules were being chosen.

Result

These days genuinely are better than average days: holding the market only on them beat holding it for the same number of days chosen at random, with a p-value of 0.025 - inside the conventional 0.05 threshold, p being the probability of a gap at least this large arising if the calendar made no difference. But the rule did not beat simply buying and holding the index, because sitting in cash through the rest of a rising 28 years costs more than the good days are worth, and it did not survive the correction for having tested 18 rules at once. The double-exposure version, which removes the cash-drag objection entirely, was not retained either.

What this does not show

"Better than average days" is not "beats owning the index", and only the second statement matters to a private investor - so this should not be read as evidence against the underlying pattern, which our own test in fact detected. The rule also switches in and out constantly, which makes it the rule in the round most sensitive to our assumed 10 basis points per switch.

The January BarometerDid not replicate
Source

Stock Trader's Almanac 2025 (Hirsch & Mistal); devised by Yale Hirsch in 1972

The claim

As January goes, so goes the year: a positive January foretells a positive February-to-December.

What we ran

If the index's January return is positive, hold the market from February through December; otherwise hold cash for the rest of the year. No lag is needed, because the signal is fully known at the January close. Same index and assumptions as the other timing rules: the 500 largest US companies by market value, excluding financials and utilities, weighted by company size, 10 basis points (0.10% of the amount traded) per switch, and cash credited at 0%, 2% and 4% a year. Judged on the same three tests fixed in advance - beat buying and holding the same index, survive the correction for having tested 18 rules at once, and be positive in both halves of the record.

Window

Daily, 1998-2026. The two halves used for the both-halves bar split at 1 January 2013.

Result

Failed the bars set in advance. The per-rule figures were displayed on screen and never saved, so what survives here is the round-level outcome: none of the 21 rules cleared all three bars, and every timing rule was negative over 2013-2026.

What this does not show

One observation a year over 28 years, and the rule is invested in most of them, so nearly all of its measured performance comes from the handful of years with a negative January - a sample far too small to say anything about the Barometer's own record. A defect in our own test also belongs here. For conditional rules like this one, the random-schedule check was documented as shuffling the SIGNAL across years, which is the version that would actually ask whether January carries information. As the test was written it shuffles the in-market days inside each calendar year regardless of the type of rule, so that check was never applied to this rule in the intended form.

The January EffectDid not replicate
Source

Stock Trader's Almanac 2025 (Hirsch & Mistal)

The claim

Small companies outperform large ones from mid-December into February.

What we ran

Universe: the 2,500 largest US companies by market value as it stood on the day, excluding financials and utilities and excluding any company whose most recent accounts were too old to use. Buy the companies ranked 1,001 to 2,500 by market value - the smaller half of that list - putting the same amount into each, on the first trading day on or after 15 December, and hold to the first trading day on or after 15 February. The basket was measured against TWO benchmarks over the identical window: the largest 500 companies weighted by company size, and the same 500 with the same amount in each. The second benchmark is the one that decides the question, and it exists precisely so that the return to holding smaller companies in equal amounts is not credited to a calendar effect. Significance came from 4,000 resamplings of the individual years.

Window

One observation a year: each mid-December to mid-February window from 1998/99 through 2025/26, about 28 in all.

Result

Ahead of the size-weighted 500 by 2.18% a year, with a p-value of 0.049 - just inside the conventional 0.05 threshold, p being the probability of a gap at least this large arising by chance if there were no effect. But against the fairer benchmark, the same 500 companies held in equal amounts, the basket was behind by a median of 0.79% a year and beat it in only 46% of years. On that benchmark the effect is absent: most of what appears against the size-weighted comparison is the basket's tilt towards smaller companies and towards equal position sizes, not a January seasonal.

What this does not show

Our "small" companies are ranks 1,001 to 2,500 of the largest US listings - mid caps, not the micro caps where the January effect is documented - and the universe excludes financials. This is a test of a large-cap-adjacent version of the claim on 28 annual observations, and it does not speak to the microcap evidence the effect was built on.

The low-volatility anomaly (and betting against beta)Not measurable here
Source

Ang, Hodrick, Xing & Zhang (2006) on idiosyncratic volatility; Blitz & van Vliet (2007) on low volatility; Frazzini & Pedersen (2014) on betting against beta. Our own tests measure volatility, market-independent volatility and beta as three separate signals, and do not follow the construction in any one of those papers.

The claim

Low-volatility and low-beta stocks - those whose prices move least, and those that move least with the market - earn as much as or more than high-volatility, high-beta ones, so risk is not rewarded inside the equity market.

What we ran

Three separate signals, each with the direction declared in advance so that the lower-risk companies were expected to win: how much a share's price moved around over the past 252 trading days, about a year; the part of that movement not explained by the market's own moves; and beta, how strongly the share moves with the market, measured over the same 252 days. Companies sorted into five equal groups on each signal, held one month, drawn from the 1,000 largest US companies by market value excluding financials and utilities, before trading costs. Separately, the same idea was applied not as a standalone signal but as a filter on Purple Turtle's own ranking, the cash-flow and profitability score behind the published books: take twice as many companies as the book needs, then keep the calmer half - either the least volatile or the lowest-beta - at five different book sizes, rebalanced twice a year and after costs.

Window

Monthly, 1998-2026 - 331 monthly observations for the three signals. Reported separately for 1998-2014 and for 2015-2026, the later stretch having been set aside before testing began and not examined while factors were being chosen.

Result

All three signals are flat to mildly opposite to the published direction, and not one of them is distinguishable from zero. Taking the gap between the extreme fifths with the same amount in each company and then weighted by company size: price volatility -0.48% and -3.88% a year (t-statistic -0.09, p-value 0.931); the market-independent part of volatility -0.09% and -4.04% (t-statistic -0.02, p-value 0.986); beta -1.70% and -3.79% (t-statistic -0.29, p-value 0.775). Every one of those measurements is a small fraction of its own margin of error, against a conventional 0.05 threshold, so none can be called a result in either direction. Split across the record, the equal-amount figures ran +1.25% then -2.91%, +1.32% then -2.07%, and -0.08% then -3.97%. The portfolio-level filter behaved differently and is worth reporting separately: keeping the calmer half of a wider pool did cut realised volatility substantially, but it turned a positive return over the index on the held-back years into a negative one. Choosing calmer COMPANIES removed what our own ranking was finding; dividing position sizes by volatility - the same companies, smaller positions in the jumpy ones - did not.

What this does not show

The published low-volatility result is a RISK-ADJUSTED one, and often a leverage-constrained one: the classic claim is a better return per unit of risk, not a higher raw return, and the betting-against-beta version is a market-neutral portfolio that buys low-beta stocks with borrowed money and sells high-beta ones short. We measured raw return gaps between five groups, with no borrowing, no short leg and no risk adjustment, so a flat result here is fully compatible with the published claim. The 2015-2026 half was also led by high-beta mega caps, which loads the dice against any low-risk tilt over exactly the years we held back.

Return seasonality (same-calendar-month persistence)Did not replicate
Source

Keloharju, Linnainmaa & Nyberg (2016), Journal of Finance, "Return Seasonalities"

The claim

Stocks with high historical returns in a particular calendar month keep outperforming in that same month - reportedly worth around 13% a year.

What we ran

For each company at each formation month, its average return in that same calendar month over the previous 20 years, requiring at least 8 such observations and lagged by one period so that the current month is not part of its own signal. Companies sorted into five equal groups on that average, held one month, drawn from the 1,000 largest US companies by market value excluding financials and utilities, before trading costs.

Window

The signal consumes at least 8 years of same-month history, so this test rests on 246 monthly observations instead of 342, covering roughly 2006-2026. Results are still reported separately for the years before 2015 and for 2015 onward, the later stretch having been set aside before testing began.

Result

Did not replicate here. The gap between the extreme fifths was 1.36% a year with the same amount in each company and 1.58% weighted by company size, against a published claim of roughly 13% a year, and under one times its own margin of error (t-statistic 0.89, p-value 0.374 against a conventional 0.05 threshold) - so it is not distinguishable from zero. It ran +2.58% before 2015 and +0.41% from 2015 on.

What this does not show

The paper estimates its result across the full breadth of the US market, small companies included, and across several seasonality horizons; ours is a single horizon on the 1,000 largest US companies excluding financials and utilities, with only about 20 usable years, because our price history starts in 1998 and the signal itself consumes the first eight. This is a thin test on the wrong universe, not a refutation.

The Santa Claus RallyDid not replicate
Source

Stock Trader's Almanac 2025 (Hirsch & Mistal); devised by Yale Hirsch in 1972

The claim

The last five trading days of December and the first two of January are reliably positive, and their absence warns of a difficult year ahead.

What we ran

Hold the daily index only on the last five trading days of December and the first two trading days of January, and hold cash for the rest of the year. Same index and assumptions as the other timing rules: the 500 largest US companies by market value, excluding financials and utilities, weighted by company size, 10 basis points (0.10% of the amount traded) charged on each switch, and cash credited at 0%, 2% and 4% a year. Judged on the same three tests fixed in advance as every other rule in the round: beat buying and holding the same index over the same window, survive the correction for having tested 18 rules at once, and be positive in both halves of the record. The same seven-day window also supplies one of the three inputs to the Almanac's January Trifecta rule, which was tested separately.

Window

Daily, 1998-2026. The two halves used for the both-halves bar split at 1 January 2013.

Result

Failed the bars set in advance. The per-rule figures were displayed on screen as the test ran and never saved, so what survives here is the round-level outcome: none of the 21 Almanac rules cleared all three bars, and every one of the timing rules came out negative over 2013-2026, which fails the both-halves bar on its own.

What this does not show

A seven-day window yields one observation a year, so 28 years is far too few to separate this from noise in either direction; the test can neither confirm the pattern nor contradict it. Our version also only asks whether those seven days are a good way to own the market. We did not test the Almanac's second and more interesting claim, that a missing rally is a warning about the year that follows.

September weaknessDid not replicate
Source

Stock Trader's Almanac 2025 (Hirsch & Mistal)

The claim

September is the worst month of the year for US equities, and can be avoided or sold short.

What we ran

Five ways of varying exposure to the daily index instead of switching all the way to cash: hold no equities, or hold a short position, for the first three trading days of September; hold half the normal amount, or a short position, for the whole month; and hold no equities from the 11th trading day of September onward. Exposure is reset daily, so the compounding path is inside the numbers and not assumed away. Borrowing costs 3% a year, with 1.5% and 5% also tested; any borrowed position carries a further 0.5% a year of friction; and 5 basis points (0.05%) is charged on every unit of exposure changed. Each of the five is scored against CONSTANT exposure at the same average level, so that only the gap over that control counts as evidence the timing added something - simply holding more of a rising market beats holding less of it on almost any schedule, and this control removes that.

Window

Daily, 1998-2026. The overlays were selected on 1998-2014 and then measured on 2015-2026, which was set aside before testing began and not examined during selection.

Result

No September overlay was kept. Across the 22 strategies derived in that round - exposure overlays, September shorts and combinations of signals - how well a strategy ranked on 1998-2014 told us nothing about how it ranked on the years held back. A caveat about our own bookkeeping applies to everything that follows, and it governs how much weight these numbers can bear. The project research log records: essentially no relationship between the two rankings (a rank correlation of -0.05); an average shortfall of 0.47 percentage points a year against the exposure-matched control on the held-back years; 9 of the 22 strategies positive there; September's own average return at -1.11% and then -0.82% across the two halves of the record against +0.90% for other months; and a median September of +0.36%. The log also records that September's average looks unlikely to be chance on its own, with a p-value of 0.021, but not once you allow for having singled out the worst of twelve months, which moves it to 0.236 - nowhere near the conventional 0.05 threshold. None of those summary figures is produced by the code as it stands, which prints only a table of the 14 best results on the held-back years. They should therefore be read as research-log figures, not as reproducible output.

What this does not show

A negative average alongside a positive median means the pattern is a fat left tail - 2001, 2008, 2022 - so this is a statement about when crashes have happened, not about a typical September. It behaves like insurance, not an edge, and our test cannot say whether that insurance is fairly priced. It also does not address the Almanac's claim on its own and far longer sample.