Do Longer STT Context Windows Improve Recent Returns?
We group 351 active STT models by 32, 64, and 128-bar context windows, then compare 30-day return distributions, position-level portfolio curves, and correlations.
> Research status: This is an observational report on 351 active STT/current runs queried on 2026-08-11. Every August 11 row carried a zero next-day return because the bar was still incomplete, so the report uses completed daily data from July 12 through August 10. It does not forecast returns, establish a causal window effect, or report live-account PnL.
The question
Does a 32, 64, or 128-bar context window change STT performance over the most recent 30 completed days?
The question has two separate measures:
1. How does the 30-day return distribution change across individual models? 2. What happens after we average target positions by symbol and window, apply the one-day execution lag, and charge costs once to the netted turnover?
The first measure describes a model population. The second describes the position-level portfolio contract used for an operating ensemble. We keep them separate.
Short answer
- The 30-day position-level returns for 32, 64, and 128 bars were +0.25%, +0.22%, and +0.27%. Their endpoints were close.
- Individual 128-bar models had the highest median return (+0.15%) and the highest positive-model share (55.9%). The same window portfolio also had the highest annualized volatility (1.80%) and the deepest drawdown (-0.40%).
- The 64-bar portfolio finished with the smallest return, but it also had the lowest annualized volatility (1.04%) and drawdown (-0.32%). A 30-day annualized Sharpe ratio is too unstable to choose a winner.
- Spearman correlation between
log2(window)and 30-day return was -0.01. Correlation between window length and total turnover was -0.41. Longer windows changed less often in this cohort; they did not show a simple return advantage.
Snapshot scope
| Item | Value |
|---|---|
| Query date | 2026-08-11; the last completed bar was 2026-08-10 |
| Cohort | 351 active STT current runs across 13 symbols |
| Window metadata | windowSize from each artifact sidecar; 351/351 mapped |
| Models by window | 32 bars: 126 / 64 bars: 114 / 128 bars: 111 |
| Evaluation range | 2026-07-12 through 2026-08-10, 30 calendar days |
| Warm-up | July 11 target positions supply the first realized exposure and turnover |
| Excluded rows | 351 incomplete zero-return rows dated 2026-08-11 |
A common market path for model and portfolio comparisons
Some model rows recorded different asset_return values for the same symbol and date because their backfills came from different data versions. Sixty-four of 390 symbol-date observations (16.4%) had more than one value. This report uses the modal value among active STT rows as the canonical market return, so every window faces the same market path.
We recalculate individual-model distributions and window portfolios under the same timing and cost contract:
ensemble_target[t] = mean(target_position[t], same symbol + same window)
realized_position[t] = ensemble_target[t - 1]
turnover[t] = abs(ensemble_target[t] - ensemble_target[t - 1])
net_log_return[t] = realized_position[t] * log(1 + canonical_asset_return[t])
- turnover[t] * execution_unit_cost[t]
position_return[t] = exp(net_log_return[t]) - 1
portfolio_return[t] = mean(symbol_position_return[t], 13 symbols)execution_unit_cost uses each current run's fee and slippage settings with the current build_strategy_curve volatility-dependent slippage formula. The daily table does not retain liquidity stress, so the calculation fixes it at the neutral value of 1.0.
The dashboard below renders the verified aggregate snapshot in the browser. We do not publish raw model rows, CSV files, or chart image artifacts.
분석 코호트
351개 active STT
13 symbols · window metadata 100% 매핑
완료된 관측 구간
30일
2026-07-12 → 2026-08-10
해석 원칙
관찰된 30일 성과이며 미래 수익 예측이 아닙니다.
32 bars
n=126+0.25%
position-level 누적 수익률
- 모델 중앙값
- +0.09%
- 양수 비율
- 53.97%
- MDD
- -0.33%
- 30D turnover
- 1.59x
64 bars
n=114+0.22%
position-level 누적 수익률
- 모델 중앙값
- 0.00%
- 양수 비율
- 48.25%
- MDD
- -0.32%
- 30D turnover
- 1.22x
128 bars
n=111+0.27%
position-level 누적 수익률
- 모델 중앙값
- +0.15%
- 양수 비율
- 55.86%
- MDD
- -0.40%
- 30D turnover
- 1.15x
30일 모델 수익률 분포
개별 target position을 같은 canonical 시장 경로·비용 계약으로 다시 계산했습니다. 막대는 각 수익률 구간에 속한 모델 수입니다.
window별 position-level 누적수익률
같은 symbol·날짜 안에서 target position을 먼저 평균하고, 전일 target을 실제 노출로 적용했습니다. 비용은 netting된 turnover에 한 번만 반영했습니다.
backbone × window
셀은 모델 30일 수익률 중앙값이며, 작은 n은 일반화 근거가 아닙니다.
-0.08%
n=33 · +45.45%
+0.77%
n=33 · +66.67%
+0.14%
n=34 · +61.76%
+0.77%
n=43 · +65.12%
-0.36%
n=37 · +43.24%
+0.44%
n=36 · +52.78%
+0.06%
n=9 · +55.56%
0.00%
n=7 · +14.29%
0.00%
n=5 · +20.00%
-0.01%
n=41 · +48.78%
-0.05%
n=37 · +43.24%
+0.17%
n=36 · +58.33%
변수 간 상관관계
Spearman 순위상관입니다. 색은 방향과 크기, 숫자는 상관계수입니다.
window만의 효과로 읽으면 안 되는 이유
active 코호트에는 서로 다른 train fraction 세대가 섞여 있습니다. 아래는 같은 window 안에서 train fraction별 모델 평균 수익률을 다시 본 값입니다.
| Train fraction | 32 bars | 64 bars | 128 bars |
|---|---|---|---|
| 10.00% | +0.63%n=44 | -0.04%n=34 | +0.29%n=27 |
| 30.00% | +0.00%n=82 | +0.34%n=80 | +0.15%n=84 |
수익률 분포는 model-level 결과, 누적 곡선은 window별 position-level ensemble 결과입니다. 서로 다른 질문에 답하므로 한 숫자로 섞어 해석하지 않습니다. 기준일 2026-08-10.
Reading the charts
128 bars followed a different path, not a decisive better endpoint
The 128-bar group led the individual-model median at +0.15% and the positive-model share at 55.9%. Portfolio endpoints still clustered tightly: +0.25% for 32 bars, +0.22% for 64 bars, and +0.27% for 128 bars.
The 128-bar portfolio fell to about -0.35% during mid-July before recovering in early August. It finished with a -0.40% maximum drawdown and 1.80% annualized volatility. Its endpoint ranked first, but the route was not smoother.
The 64-bar portfolio reached higher interim values, then closed at +0.22%. Its 1.04% volatility and -0.32% drawdown were the lowest of the three. The sample is too short to turn that observation into a risk-adjusted ranking.
Window length tracked turnover, not recent return
The correlation matrix gives ρ = -0.01 for log2(window) and 30-day return. Window length alone did not explain the cross-sectional return order in this active snapshot.
The relationship with turnover was clearer. The window-level total turnover fell from 1.59x at 32 bars to 1.22x at 64 bars and 1.15x at 128 bars, while the Spearman correlation was -0.41. Longer input histories coincided with less frequent signal changes.
Average absolute position and annualized volatility had a +0.76 correlation. Return and average absolute position had a +0.02 correlation. Larger exposures raised observed volatility here; they did not explain return differences.
Backbone and training generation confound a window-only claim
The backbone-by-window cells do not move in one direction. FASCL had median returns of -0.08%, +0.77%, and +0.14% across 32, 64, and 128 bars. Franceschi showed +0.77%, -0.36%, and +0.44%. TS2Vec reached +0.17% at 128 bars while its 32- and 64-bar medians were close to zero or negative.
The active cohort also mixes 105 models trained with a 10% train fraction and 246 trained with a 30% train fraction. Within 32 bars, the 10% generation averaged +0.63% and the 30% generation was nearly flat. Within 64 bars, the direction reversed.
These models do not isolate window size. Backbone, training fraction, and model generation all move with the window groups.
The 5Y gate score and this 30-day slice answer different questions
Spearman correlation between the 5Y gate Sharpe and recent 30-day return was +0.03. That does not invalidate the gate. It says that, across this active STT cohort and this short cross-section, the historical gate score did not rank the next 30 days of returns.
The gate summarizes a much longer selection period. This report observes one completed month. They should not be treated as interchangeable metrics.
Observed, unproven, next
- Observed: all three window portfolios posted positive 30-day returns between +0.22% and +0.27%.
- Observed: 128 bars led the model-level median and positive share, while its portfolio had the highest volatility and deepest drawdown.
- Observed: longer windows coincided with lower turnover in this snapshot.
- Unproven: increasing context length causes better return.
- Unproven: these 30-day results will persist over later 30-day or 90-day windows, or in live-account PnL.
- Next: fix matched cohorts by backbone, symbol, train fraction, and seed before training, then repeat this aggregation using only completed forward bars.
Limits
- The report applies the current active model set to the preceding 30 days, so selection bias remains.
- Models from the same family, symbol, and training generation are not independent observations.
- Funding, unfilled orders, latency, account leverage, and account-level risk limits are outside this calculation.
- The canonical modal market return makes the window comparison consistent; it does not reconstruct the original data version of every historical row.
- STT's random sliding-window train/test split and full-period gate are not independent out-of-sample evidence. This report does not turn one month of observations into that evidence.
This is a research record, not investment advice.