Methodology

The three baseline factors are a replication of the exact methodology described by Liu et al. (2022) in Section II. Similar to Liu et al. (2022), we use CoinMarketCap as source for the daily close price, volume, and market capitalisation (USD) of all coins. CoinMarketCap collects data from all major exchanges and aggregates them.

The data is retrieved via the crypto2 R package, which provides a survivorship-bias-free cross-section of all crypto assets (including delisted/inactive coins). Weekly factors are constructed from the daily data using the same breakpoints and weighting schemes as Liu et al. (2022).

Our sample period ranges from January 2014 to present (the paper covers January 2014 to July 2020). Each calendar year is divided into exactly 52 weeks (Liu et al. 2022, Section I). The first week of the year is January 1–7, weeks 2–51 are 7 days each, and week 52 is the remaining 8 days (9 days in leap years).

Before constructing the factors, a similar filtering logic as Liu et al. (2022) is applied. Each coin-week must satisfy:

Liu et al. (2022) do not filter for minimum age or stablecoins. However, we provide these options as a specifications of the downloadable factor series.

CMKT — Crypto Market Factor

The market factor is the value-weighted return of all eligible coins, minus the risk-free rate. The risk-free rate is proxied by the 1-month T-bill rate, which is negligible at weekly frequency and set to 0.

\[\text{CMKT}_t = \sum_{i} w_{i,t}\, r_{i,t} \;-\; R_{f,t}\]

where \(w_{i,t}\) = lagged market cap weight and \(R_{f,t}\) = 1-month T-bill rate (negligible at weekly frequency; set to 0).

CSMB — Crypto Size Factor

To construct the size factor, each week, eligible coins are sorted by lagged market capitalisation into three groups using 30/40/30 breakpoints. The groups are defined as:

  • Small = bottom 30% by market cap
  • Neutral = middle 40% (excluded from the factor)
  • Big = top 30% by market cap

In each group , the return is computed as the value-weighted return of all coins in that group. The size factor is then the difference between the small and big groups:

\[\text{CSMB}_t = R_{\text{Small},t} - R_{\text{Big},t}\]

CMOM — Crypto Momentum Factor

CMOM is built with a Fama-French 2×3 independent double sort on size and momentum. Each week the coins are sorted into two size portfolios (Small / Big; 50/50) and three momentum portfolios (Loser / Neutral / Winner; 30/40/30). The momentum signal is the 3-week cumulative return (\(r_{3,0}\)) ending at the formation date:

\[\text{mom}_{i,t} = \frac{P_{i,t-1}}{P_{i,t-4}} - 1\]

where \(P_{i,t}\) = end-of-week price. In our weekly panel, this is lag(price_eow, 1) / lag(price_eow, 4) - 1.

As a result four corner portfolios are received:

Loser (L) Winner (H)
Small (S) SL SH
Big (B) BL BH

Within each cell the returns are value-weighted and the momentum factor is then the average of the two winner portfolios minus the average of the two loser portfolios:

\[\text{CMOM}_t = \frac{1}{2}(R_{SH,t} + R_{BH,t}) - \frac{1}{2}(R_{SL,t} + R_{BL,t})\]

Summary of Liu et al. Parameters

Parameter Value
Weighting Value-weighted (lagged market cap)
Size breakpoints 30% / 40% / 30%
Momentum breakpoints 30% / 40% / 30%
Momentum signal \(r_{3,0}\) (3-week return, no skip)
CMOM structure 2×3 Fama-French independent double sort
Market cap filter \(\geq\) $1,000,000 (lagged)
Risk-free rate 1-month T-bill (set to 0 at weekly frequency)
Calendar Sharp year (52 weeks, Jan 1–7 = week 1)
Stablecoin exclusion None
Winsorisation None documented

Data Quality

CoinMarketCap, does not apply exchange-level outlier detection. We therefore apply two filters to further ensure reliability of the data:

Return Cap

Individual coin weekly returns are capped at 9,900% (100×). This handles CMC price/supply extrem outliers (e.g., SquidGrow: 2,597,962% weekly return). It affects ~870 coin-weeks. No legitimate crypto asset sustains 100× in a single week.

Implied Supply Filter

Coin-weeks where market_cap / price exceeds \(10^{16}\) tokens are excluded. This catches erroneous circulating supply data (e.g., INNBCL: price $0.00000001 with $126B market cap, implying \(1.3 \times 10^{19}\) tokens). Affects ~210 coin-weeks. Legitimate high-supply coins like Shiba Inu (\(\sim 5 \times 10^{14}\) tokens) are not affected.

Early Sample Period

Before July 2014, the CoinMarketCap universe contains fewer than 30 eligible coins. Factor estimates are noisy and dominated by Bitcoin, Litecoin, and a handful of other early coins. We recommend starting analyses from July 2014. Plots on this website use this cutoff.


The specification multiverse

Rather than report one “preferred” set of numbers, we construct the factors across a multiverse of defensible specification choices and report the distribution of results. The full results table (one row per specification, with factor means and Fama-MacBeth pricing statistics) is on the Data page; the findings are summarised in the Results.

Specification axes

Axis Options Note
Data source CoinMarketCap, CoinGecko treated as a specification choice, not a fixed input
Breakpoint universe All coins, Top-100, $100M floor the headline axis (see below)
Evaluation universe Full cross-section, Investable which assets the model is asked to price
Gap-handling Consecutive-week, Naive see below
Investable-momentum On, Off see below
Weighting Value-weighted, Equal-weighted
Size breakpoints 2 (median), 3, 5 (quintile), 10 (decile)
Momentum lookback 1, 2, 3, 4 weeks
Calendar Sharp year, Mon … Sun (7 weekday starts) calendar/frequency are alignment boundaries — not cross-priced
Exclusions None, Stablecoins, +Wrapped/derivatives
Delisting returns Off, On

The headline weekly grid is ~6,900 internally-consistent “worlds” (each builds both the factors and the test assets with the same choices).

Breakpoint universe (the NYSE-breakpoint analog)

Crypto has no listing threshold — anyone can mint a token and list it on a cryptocurrency exchange — so if size/momentum breakpoints are computed over all coins, the bottom deciles are pure microcaps and the sort is dominated by economically irrelevant assets. Additionally, these microcaps are hard-to-arbitrage and illiquid, resulting in the factor being not investable under real world constraints. The multiverse therefore includes a breakpoint universe axis, which defines the set of coins used to compute the size/momentum breakpoints.

Following the Fama-French convention of computing breakpoints on NYSE stocks, we compute the cut points on an investable reference set (top-N by market cap, or a market-cap floor) and then assign all coins. This prices the full cross-section (no cherry-picking) with economically meaningful, snapshot-stable breakpoints. It is the single most consequential choice — see Results.

Gap-handling (consecutive-week returns)

A coin with a data gap, naively, produces a multi-week return mislabelled as one week (price / previous-available-price). These spurious extremes land in the size/momentum tails. The consecutive-week rule uses a return only when the prior calendar week actually exists.

Investable-momentum

Momentum is measured only over weeks in which the coin was continuously investable (≥ $1M throughout the lookback). Otherwise a coin that just crossed the size threshold on a one-off pump enters the winner portfolio with about-to-reverse “momentum”.


Data Pipeline

  1. Retrieval: Daily cryptocurrency snapshots from CoinMarketCap via crypto2::crypto_listings(). Survivorship-bias free.
  2. Processing: Returns computed from end-of-week prices. Return cap (9,900%) and implied supply filter applied. Aggregated to weekly (sharp-year or Monday-Monday) and monthly panels.
  3. Liu exact factors: compute_factors_liu() with hardcoded 30/40/30 breakpoints → factors_liu.parquet
  4. Variant factors: compute_factors() with parametric breakpoints across all 576 combinations
  5. Portfolios: compute_portfolios() — decile sorts on size and momentum
  6. Export: All files uploaded to Hugging Face as CSV/Parquet.

Citation

If you use these data, please cite Stoeckl & Pukrop (2026) and Liu et al. (2022).

@article{liu2022common,
  author  = {Liu, Yukun and Tsyvinski, Aleh and Wu, Xi},
  title   = {Common Risk Factors in Cryptocurrency},
  journal = {The Journal of Finance},
  volume  = {77},
  number  = {2},
  pages   = {1133--1177},
  year    = {2022},
  doi     = {10.1111/jofi.13119}
}


@misc{stoeckl2026opencrypto,
  author      = {Stoeckl, Sebastian and Pukrop, Moritz},
  title       = {Open Crypto Pricing},
  year        = {2026},
  institution = {University of Liechtenstein},
  url         = {https://huggingface.co/datasets/sstoeckl/opencryptoassetpricing}
}

References

Liu, Y., Tsyvinski, A., & Wu, X. (2022). Common risk factors in cryptocurrency. The Journal of Finance, 77(2), 1133–1177. https://doi.org/10.1111/jofi.13119
Stoeckl, S., & Pukrop, M. (2026). Open crypto pricing. University of Liechtenstein. https://huggingface.co/datasets/sstoeckl/opencryptoassetpricing