TTT Ouroboros
arXiv 2610.05076 · cs.CL · October 20262026 年 10 月

Self-Generated Feedback Destabilizes Test-Time Training

自生成反馈让测试时训练失稳

Self-Generated Feedback Destabilizes Test-Time Training

A causal decomposition of long-horizon adaptation长时程适应的因果分解

Cheng LuoBing Li*Bernard Ghanem
King Abdullah University of Science and Technology (KAUST) · *corresponding author阿卜杜拉国王科技大学(KAUST)· *通讯作者
tap the snake to redraw it点一下蛇,重新画一遍

Some language models keep learning while they read. That is the promise of test-time training. Now give such a model a long stream of its own writing to learn from. Over 128K tokens it slowly gets worse at reading everyone else's.

有些语言模型会边读边学,这正是测试时训练的卖点。现在,让这样一个模型拿自己写出的文字来学习。读完 128K 个 token 之后,它读懂别人文字的能力在一点点变差。

This page tells the story of how we tracked down why, one controlled comparison at a time, and of a simple rule that stops it: check an update on real text before you keep it.

这一页讲的是我们如何用一组组对照实验,一步步找出原因;以及一条能阻止它的简单规则:先在真实文本上检验一次更新,再决定要不要留下它。

The snake eating its tail is the whole paper in one picture.咬住自己尾巴的蛇:整篇论文就是这一张图。
+3.01

nats worse at predicting real text after 128K tokens of self-writing (125M model)

125M 模型自我写入 128K token 后,预测真实文本的损失多出 3.01 nats

98%

of that damage disappears when a frozen copy does the writing instead

换成一个冻结的副本来写,这部分损害几乎全部消失

.07

nats left once Settlement checks each update on independent real text

用 Settlement 在独立真实文本上检验每次更新后,只剩下这么多 nats

Watch the paper观看论文视频
TTT Ouroboros video thumbnail

Watch on YouTube ↗在 YouTube 观看 ↗

Chapter 1第 1 章

A model that learns while it reads

一个边读边学的模型

Test-time training (TTT) lets part of a model, its fast weights Wt, keep changing during inference. Information can then live in the weights after the text has scrolled out of the attention window. That is attractive for long documents and for agents that run for hours.

测试时训练(TTT)让模型的一部分参数,也就是快权重 Wt,在推理时继续更新。这样,文本即使已经滑出注意力窗口,信息仍能留在权重里。对长文档和长时间运行的智能体来说,这很有吸引力。

Two words will carry the rest of the story.

接下来的故事里,有两个词会反复出现。

Reading读(reading)

The text sits in attention and shapes the next predictions. Nothing permanent changes.

文本留在注意力里,影响接下来的预测。没有任何东西被永久改变。

Writing写(writing)

The gradient update learned from the text is kept. The weights themselves change.

从这段文本学到的梯度更新被保留下来,权重本身变了。

Writing is not the villain. On real books, it helps:

写入本身不是坏人。在真实的书上,它是有帮助的:

Keeping updates from real text improves prediction保留真实文本的更新,预测会变好
benefit B = NLL without the update − NLL with it, in nats · higher is better收益 B = 不更新时的 NLL − 更新后的 NLL(nats)· 越高越好
On eleven teacher-forced books the input is fixed, so the .034-nat gain comes from learning itself. Qwen3-4B rows show the range of two adapted states.
在 11 本教师强制输入的书上,输入是固定的,所以 .034 nats 的收益只能来自学习本身。Qwen3-4B 两行显示的是两个适应状态的范围。
Remember this. The same update rules that help here will fail in chapter 3.记住这一点。同样的更新规则,到第 3 章就会失灵。
Chapter 2第 2 章

The model writes its own homework

模型给自己出作业

Now let the model supply its own training data. The current weights Wt write a 1024-token chunk xt. The model learns from xt and becomes Wt+1. Then Wt+1 writes the next chunk. We call this the Closed Loop.

现在让模型自己提供训练数据。当前权重 Wt 写出一段 1024 token 的文本 xt,模型从 xt 中学习,变成 Wt+1;然后由 Wt+1 写下一段。我们把它叫做闭环(Closed Loop)。

To measure what the loop costs, we need a fair control. Writes Off reads exactly the same kind of generated text, keeping it in attention, but throws each update away.

要衡量这个闭环的代价,需要一个公平的对照组。关闭写入(Writes Off)读的是同样的生成文本,也把它们留在注意力里,但每次更新都直接丢掉。

Lower loss on xt only shows the update fits xt. It says nothing about other text.xt 上的损失变低,只说明这次更新拟合了 xt,不代表它对别的文本更好。
Closed Loop闭环 · the thing we study· 研究对象
Writes Off关闭写入 · the control· 对照组
H = Dclosed − DoffD is the change in real-text NLL from the first to the last check. H > 0 means keeping generated-text updates made things worse.D 是真实文本 NLL 从第一次检查到最后一次检查的变化。H > 0 表示保留生成文本的更新让结果变差。

Every check uses independent human-written passages from later, disjoint parts of the same PG-19 book. After each check, weights and attention are restored, so checking never changes the stream.

每次检查都用独立的人类文本:来自同一本 PG-19 书中靠后、互不重叠的段落。每次检查完都会恢复权重和注意力状态,所以检查本身不会改变数据流。

Chapter 3第 3 章

Nothing happens. Then everything does.

起初风平浪静,然后突然崩坏

We ran both policies for 128K tokens: an 8K real-text prefix, then 105 generated chunks. In the first few chunks the two look almost the same. Inside the 8K+8K window used by earlier evaluations, you would see nothing wrong. Only after many more self-writes does the Closed Loop pull away.

我们让两种策略各跑 128K token:先是 8K 的真实文本前缀,然后是 105 段生成文本。最初几段里,两者几乎一样。在以往评测采用的 8K+8K 窗口内,你看不出任何问题。只有经过更多次自我写入之后,闭环才开始明显偏离。

A short benchmark would have called this method safe.如果只跑一个短基准,这个方法会被判定为安全。
Extra damage from keeping generated-text updates保留生成文本更新带来的额外损害
H in nats, with 95% intervals over books · same six books, five seeds, 128K at every scaleH(nats),95% 区间按书重采样 · 每个规模都用同样 6 本书、5 个种子、128K
Every TTT-E2E interval excludes zero. The size of the effect is not monotone in model size, so read this as the same direction replicated at three scales, not a scaling law. The bottom rows drop TTT entirely: plain Adam on Qwen3-4B's own feed-forward weights fails at learning rate 10−4, the same setting that gains .168 nats on real text.
所有 TTT-E2E 的区间都不包含零。效应大小并不随模型规模单调变化,所以应理解为同一方向在三个规模上复现,而不是缩放规律。最下面两行完全去掉了 TTT:直接用 Adam 更新 Qwen3-4B 自带的前馈层权重,学习率 10−4 时同样失败;而同一设置在真实文本上能带来 .168 nats 的收益。
Where each policy starts and ends每种策略的起点和终点
real-text NLL · lower is better真实文本 NLL · 越低越好
Closed Loop闭环Fixed Generation固定生成Writes Off关闭写入
All three start from exactly the same loss. Only the endpoints are drawn; the full trajectory is the paper's Figure 1a. Fixed Generation shows up properly in chapter 4.
三者起点完全相同。这里只画了首尾两点,完整轨迹见论文图 1a。固定生成会在第 4 章正式登场。

By the end, the Closed Loop's generated text is almost pure repetition: 99% of its 4-grams repeat (repeated-4 = .9902), against .0217 for Writes Off.

到最后,闭环生成的文本几乎全是重复:99% 的 4-gram 是重复的(repeated-4 = .9902),而关闭写入只有 .0217。

Could a better decoder fix it? Not reliably. Full-support sampling halves the 128K gap to 1.452 nats, but its loss is still climbing at the end. Lower temperature, typical sampling and a repetition penalty change how fast things go wrong, not whether a given update is safe.

换个解码方式能解决吗?不太靠谱。全支撑采样把 128K 时的差距减半到 1.452 nats,但到最后损失仍在上升。降低温度、typical 采样、重复惩罚,改变的是坏得多快,而不是某一次更新是否安全。

So what exactly is going wrong?

那么,到底是哪里出了问题?

Chapter 4第 4 章

Three suspects

三个嫌疑人

There are three plausible culprits. Maybe generated text is simply bad training data. Maybe degraded text hurts while it sits in attention. Maybe learning from it stores the harm in the weights. We question each one by cutting one link at a time.

有三个看起来都说得通的嫌疑人。也许生成文本本身就是糟糕的训练数据;也许退化的文本只要待在注意力里就会造成伤害;也许从中学习会把伤害存进权重。我们每次只切断一条链路,逐个审问。

Before we start: make a guess开始之前,先猜一猜

Suppose a frozen copy W0 writes all the text, and the learner still learns from every chunk. How much of the +3.01 nats of damage goes away?

假设所有文本都由一个冻结的副本 W0 来写,学习者照样从每一段学习。+3.01 nats 的损害会消失多少?

40%

Now scroll through the investigation. The sketch follows along.现在往下滚动,开始审问。草图会跟着变化。

The scene案发现场

One model, writing and learning in a circle

一个模型,在一个圈里边写边学

Wt writes the chunk, learns from it, and the updated Wt+1 writes the next one.

Wt 写出这一段,从中学习,更新后的 Wt+1 再写下一段。

Damage after 128K tokens: +3.01 nats at 125M, +6.00 at 760M, +0.40 at 3B.
128K token 后的损害:125M 上 +3.01 nats,760M 上 +6.00,3B 上 +0.40。
Suspect 1嫌疑人 1

Is generated text just bad food?

生成文本只是“垃圾食品”吗?

Let a frozen copy W0 write every chunk. A separate learner still learns from every one of them. The learner still eats synthetic text; it just can no longer change what it will be fed next. This is Fixed Generation.

让一个冻结的副本 W0 来写每一段,另一个独立的学习者照样从每一段学习。学习者吃的仍是合成文本,只是再也无法改变自己下一顿吃什么。这就是固定生成(Fixed Generation)。

Damage left: .052 nats. That removes 98.3% of the gap (98.8% at 760M).
剩余损害:.052 nats,消除了 98.3% 的差距(760M 上为 98.8%)。
Not enough on its own单独不足以解释
Suspect 2嫌疑人 2

Does reading degraded text hurt?

读退化的文本会受伤吗?

Record a stream that has already degraded. Hand it to a fresh receiver trained on other books, and let it only read, with no updates.

把一段已经退化的生成流录下来,交给一个在其他书上训练过的全新接收者,只让它读,不做任何更新。

Reading alone adds 1.606 nats while the text sits in attention.
光是读,就在文本停留于注意力期间多出 1.606 nats。
Hurts有伤害
Suspect 3嫌疑人 3

Does writing it down make it worse?

把它写进权重会更糟吗?

Same receiver, same tokens, second pass. This time it keeps the updates. Because the text is identical, any extra harm was stored in the weights.

同一个接收者、同样的 token,第二遍这次保留更新。因为文本完全相同,多出来的伤害只能是存进了权重。

Total 3.874 nats: writing adds 2.268. At 3B the same contrast is +.193 [.104, .327].
总计 3.874 nats:写入又多出 2.268。3B 上同样的对比为 +.193 [.104, .327]。
Hurts more伤害更大
The verdict结论

The loop is what turns text bad

是闭环把文本变坏的

Reading and writing degraded text both hurt. But the text only degrades because each update feeds back into the model's own future training data. Cut that link and almost nothing is left.

读和写退化文本都有伤害。但文本之所以会退化,是因为每次更新都会反馈进模型自己未来的训练数据。切断这条链路,几乎什么都不剩。

It is not about how far the weights move, either. Fixed Generation moves them farther than Closed Loop (drift .0139 vs .0086) and does far less harm.

这也和权重移动了多远无关:固定生成比闭环移动得更远(漂移 .0139 对 .0086),伤害却小得多。

Main cause: feedback主因:反馈

The evidence, side by side证据并排放

Cut the feedback切断反馈
damage over Writes Off at 128K, nats128K 时相对关闭写入的损害(nats)
Same text, read vs. written同样的文本:只读 vs 写入
NLL above the receiver's control, nats相对接收者对照组多出的 NLL(nats)
Drift doesn't rank damage漂移幅度解释不了损害
matched 125M control · 8 books, 5 seeds匹配的 125M 对照 · 8 本书、5 个种子
Chapter 5第 5 章

One update, two faces

一次更新,两副面孔

The endpoints above mix thousands of updates. To see what a single update does, copy the model right before a generated chunk: same weights, same attention, same text. One copy keeps the update. The other throws it away. Any difference afterwards is caused by that one update.

上面的终点混合了成千上万次更新。要看清单次更新做了什么,就在生成一段文本之前把模型复制一份:同样的权重、同样的注意力、同样的文本。一份保留这次更新,另一份丢掉。之后的任何差异,都来自这一次更新。

A better fit to its own words, a worse fit to ours.对自己的话拟合得更好,对我们的话预测得更差。
after Writes Off history关闭写入历史之后after Closed Loop history闭环历史之后after Fixed Generation history固定生成历史之后
Its own chunk: keep − discard它自己那段:保留 − 丢弃
NLL change · left = fits its source betterNLL 变化 · 越靠左越拟合源文本
The next real passage: keep − discard下一段真实文本:保留 − 丢弃
NLL change · right = hurts new textNLL 变化 · 越靠右伤害越大
Every single row: better on its own chunk, worse on new real text. After a Closed Loop history the gain on its own chunk shrinks to .0354 and the cost on real text grows to .0381. 125M, 95% book intervals.
每一行都一样:在自己那段上更好,在新的真实文本上更差。经历闭环历史之后,自身那段的收益缩小到 .0354,而真实文本上的代价增加到 .0381。125M,95% 按书区间。

The cost also grows as the loop runs. Early on, keeping one update costs .006 nats. Late in a Closed Loop it costs .113, while after Writes Off or Fixed Generation it stays near zero.

而且随着闭环运行,代价越来越大。早期保留一次更新代价 .006 nats;闭环后期变成 .113,而在关闭写入或固定生成的历史之后几乎为零。

Cost of keeping one update, by when it happens保留一次更新的代价,按发生的时间
next real-text NLL, keep − discard, nats · 125M下一段真实文本 NLL,保留 − 丢弃(nats)· 125M
Closed Loop history闭环历史Writes Off history关闭写入历史Fixed Generation history固定生成历史

The gradients disagree梯度在“打架”

Before keeping an update, compare the gradient from the generated chunk with the gradient from the next real passage. The more they point in opposite directions, the more the update hurts. Over 96 paired updates per setting, the rank correlations are strong and every interval excludes zero.

在保留一次更新之前,比较生成片段的梯度和下一段真实文本的梯度。方向越相反,这次更新伤害越大。每种设置 96 对更新,秩相关都很强,所有区间都不包含零。

Model / history模型 / 历史ρ(cos, harm伤害)ρ(g⊤ΔW, harm伤害)mean cos平均 cos
125M · Closed Loop闭环−.502.704−.093
125M · Writes Off关闭写入−.788.794−.005
760M · Closed Loop闭环−.633.632−.103
760M · Writes Off关闭写入−.731.751−.004

After a Closed Loop history, generated and real gradients actively oppose each other (mean cos ≈ −.10); after Writes Off they are nearly orthogonal. This explains which updates hurt, but it is not a usable filter, since it peeks at the very passage on which harm is measured.

经历闭环历史后,生成文本与真实文本的梯度明显相反(平均 cos ≈ −.10);关闭写入之后则几乎正交。这解释了哪些更新有害,但不能直接当过滤器用,因为它偷看了用来衡量伤害的那段文本。

Chapter 6第 6 章

Most late updates are fine. A few are terrible.

大多数后期更新没事,少数是灾难

If late updates get more expensive on average, are they all bad? We replayed eight-chunk passages into identical receivers, once read-only and once read-and-write, 72 pairs at each position.

如果后期更新平均越来越贵,是不是每一个都坏?我们把 8 段长的片段回放给完全相同的接收者,一次只读、一次读并写,每个位置 72 对。

The mean rises; the median doesn't. A handful of runs carry the whole average.均值在涨,中位数不动。整个平均值是少数几条轨迹撑起来的。
Typical vs. average write cost典型 vs 平均写入代价
NLL(read + write读+写) − NLL(read only只读), nats
mean均值median中位数
Disasters: costs above 0.5 nats灾难:代价超过 0.5 nats
count out of 72 paired comparisons72 对比较中的个数

At the last position the mean is .7512 but the median is .0064. All twelve disasters come from 4 of 24 source sequences, and once a pair crosses the line it stays above it. So "late in the stream" is a warning sign, not a rule for which update to drop.

在最后一个位置,均值是 .7512,中位数却只有 .0064。全部 12 次灾难来自 24 条源序列中的 4 条,而且一旦某对越过阈值,之后就一直在阈值之上。所以“处在数据流后期”只是警示信号,不能用来决定该丢哪次更新。

Chapter 7第 7 章

Half-fixes

治标不治本

Before building something new, we tried the obvious remedies. Each one helps. None of them can tell you whether a particular update is safe to keep.

在设计新方法之前,我们先试了显而易见的补救办法。每一个都有点用,但没有一个能告诉你某一次更新是否值得保留。

Remedy 1 · Interrupt the loop with real text补救 1 · 用真实文本打断闭环

Swap some of the 105 generated chunks for real ones, still keeping every update. Try the schedules below and watch the longest stretch the model spends writing to itself.

把 105 段生成文本中的一部分换成真实文本,同时仍保留每一次更新。试试下面几种安排,观察模型连续“自说自话”的最长一段。

real chunk真实文本段self-written chunk自写段longest self-written stretch最长连续自写
longest self-written stretch最长连续自写
3
chunks in a row段连续
damage H损害 H
.0694
vs. no real text相对没有真实文本
6%
Spacing matters as much as amount. At 31%, real text spread out leaves 6% of the damage; the same 33 chunks in bursts leave about 60%. Separate eight-book long-stream suite, 125M, five seeds.
分布方式和数量一样重要。31% 时,均匀分布的真实文本只留下 6% 的损害;同样 33 段成团出现,则留下约 60%。独立的 8 本长书实验,125M,5 个种子。

Remedies 2–4 · Write less, decode differently, down-weight repeats补救 2–4 · 写得更少、换解码方式、给重复降权

α = 1

Chapter 8第 8 章

Ask before you remember

先问一句,再记下来

Every half-fix shares one blind spot: none asks whether the update actually helps beyond its own chunk. So ask directly. Keep each update pending. When independent real text q arrives, compare the current state with the candidate state on it. Keep the update only if the candidate predicts q at least as well. We call this Settlement.

所有半吊子方案都有同一个盲点:没有一个去问,这次更新在它自己那段之外是否真的有帮助。那就直接问。让每次更新先挂起,等独立的真实文本 q 到来时,在 q 上比较当前状态和候选状态,只有候选状态预测 q 至少一样好,才保留这次更新。我们把它叫做 Settlement(先验证,再提交)。

The gate never looks at where the text came from. It only looks at what the update does.这道关卡从不看文本是从哪来的,只看这次更新做了什么。
A(δ; W, q) = L(q; W) − L(q; W + δ)commit δ only if A ≥ 0只有 A ≥ 0 时才提交 δ
# Algorithm 1, condensed
q ← newest real text (or the retained bank)
S ← W
for δ in pending, in order:
    if NLL(q; S + δ) ≤ NLL(q; S):
        S ← S + δ    # judge the next δ on top
commit W ← S
# 算法 1(精简版)
q ← 最新的真实文本(或保留的文本库)
S ← W
for δ in 待定更新,按提出顺序:
    if NLL(q; S + δ) ≤ NLL(q; S):
        S ← S + δ    # 下一个 δ 在此基础上评判
提交 W ← S
The result结果
gap to Writes Off at the end of the stream, nats, 95% interval数据流末端与关闭写入的差距(nats),95% 区间
Closed Loop闭环Settlement
Almost all the damage is gone. Settlement keeps only 22 of 936 generated-text updates at 125M and 18 of 312 at 760M. Both leftover intervals include zero; that is not a claim of equivalence. It beats Closed Loop on 11 of 12 books at 125M and on all four at 760M.
损害几乎全部消失。Settlement 在 125M 上只保留了 936 次生成文本更新中的 22 次,在 760M 上保留了 312 次中的 18 次。两个剩余区间都包含零,但这不等于声称两者等价。它在 125M 的 12 本书中有 11 本优于闭环,在 760M 的 4 本书上全部优于闭环。

Check the state you will actually keep检验你真正要保留的那个状态

One subtlety. Updates that each look safe alone can be harmful together. So every candidate is judged on top of the ones already accepted, never against the original weights.

还有一个细节:单独看都安全的更新,合在一起可能有害。所以每个候选都在已接受的更新之上评判,而不是对着原始权重。

H: harm on generated streams (lower is better). B: benefit on eleven real-text books vs. no updates (higher is better). H and B come from separate streams.

H:生成文本流上的伤害(越低越好)。B:在 11 本真实文本书上相对不更新的收益(越高越好)。H 和 B 来自不同的数据流。

The check does the work, not smaller steps起作用的是检验,不是更小的步子
final NLL above Writes Off, nats · 125M, four books, three seeds相对关闭写入的最终 NLL 增量(nats)· 125M,4 本书、3 个种子
without the check不检验with Settlement加上 Settlement
And it still learns from real text它仍然能从真实文本中学习
improvement vs. rejecting every update, nats · all-real stream相对拒绝全部更新的改进(nats)· 纯真实文本流
On an all-real stream, Settlement admits 89 of 113 updates and edges out writing everything (+.0083 [.0020, .0161]). It rejects generated updates because they fail the check, not because it rejects everything.
在纯真实文本流上,Settlement 接受了 113 次更新中的 89 次,还略胜于全部写入(+.0083 [.0020, .0161])。它拒绝生成文本的更新,是因为这些更新没通过检验,而不是因为它什么都拒绝。

What if you just labeled the text?如果只是给文本打上标签呢?

A simpler rule, Source Masking, keeps updates from text labeled "real" and drops those labeled "generated". With perfect labels it is simpler and slightly better. Now corrupt the labels and drag the slider.

一个更简单的规则叫来源屏蔽(Source Masking):保留标为“真实”的文本带来的更新,丢掉标为“生成”的。标签完全正确时,它更简单,也略好一点。现在把一部分标签弄错,拖动滑块看看。

38.9%
Source Masking来源屏蔽Settlement
generated updates Masking wrongly keeps屏蔽错误保留的生成更新
26
real updates Masking wrongly drops屏蔽错误丢弃的真实更新
18
Settlement − MaskingSettlement − 屏蔽
−.0369
Settlement barely notices. It stays between 3.8196 and 3.8286 at every error rate. Even with every label wrong, it admits 120 of 120 truly real updates and 0 of 219 self-generated ones. 125M mixed stream, three seeds; lower is better.
Settlement 几乎不受影响。在所有错误率下,它都保持在 3.8196 到 3.8286 之间。即便标签全错,它也接受了全部 120 次真正来自真实文本的更新,而 219 次自我生成的更新一次也没接受。125M 混合数据流,3 个种子;越低越好。
Chapter 9第 9 章

Out in the world: agents

走进真实世界:智能体

Agents that learn from their own actions face the same loop, with a twist: every update carries the same "generated" label. A label rule can only keep them all (Closed Loop) or drop them all (Writes Off). A task score can tell them apart.

从自己的行动中学习的智能体,面对的是同一个闭环,还多一个麻烦:每次更新都带着同一个“生成”标签。基于标签的规则只能全部保留(闭环),或全部丢弃(关闭写入)。任务得分则能把它们区分开。

WebShop · exact successWebShop · 完全成功率
Qwen3-4B with LoRA · mean ± SD, five paired seedsQwen3-4B + LoRA · 均值 ± 标准差,5 个配对种子
Settlement checks validation reward every 25 episodes and keeps 10 of 180 blocks. Fixed Generation and Settlement each beat Closed Loop on every paired seed.
Settlement 每 25 个回合检查一次验证奖励,在 180 个更新块中只保留 10 个。固定生成和 Settlement 在每个配对种子上都胜过闭环。
ALFWorld · unseen household tasksALFWorld · 未见过的家务任务
Qwen3.8-27B with LoRA · % solved, mean ± SD, 3 seedsQwen3.8-27B + LoRA · 成功率 %,均值 ± 标准差,3 个种子
Two of three Closed Loop seeds solve none of the 134 unseen tasks. Settlement solves 88–94% per seed while keeping 19–33% of update blocks. Three seeds are not enough to claim a gain over Writes Off.
闭环的 3 个种子里,有 2 个在 134 个未见任务上一个都没完成。Settlement 每个种子完成 88–94%,同时保留 19–33% 的更新块。3 个种子还不足以宣称它优于关闭写入。
Epilogue尾声

What we learned

我们学到了什么

  1. 1The loop is the problem, not the text. A learner that eats synthetic text from a frozen writer stays close to Writes Off. Damage grows when the learner also writes its own future lessons.问题出在闭环,而不是文本。学习者从冻结的写手那里吃合成文本,表现仍接近关闭写入。只有当学习者也在给自己写未来的课本时,损害才会增长。
  2. 2Damage is concentrated. Four of 24 source sequences produce every disaster above 0.5 nats. The loop makes some trajectories persistently harmful, not every late update.损害高度集中。24 条源序列中的 4 条造成了所有超过 0.5 nats 的灾难。闭环让某些轨迹持续有害,而不是让每次后期更新都有害。
  3. 3Fitting its own chunk proves nothing. Neither source fit nor repetition tells you whether an update helps on new text. Real text helps most when it interrupts often.拟合好自己那段,什么也证明不了。源文本拟合度和重复度都无法说明一次更新在新文本上是否有用。真实文本最有效的用法,是频繁地打断自我写入。
  4. 4Check on independent evidence, then commit. Settlement nearly matches Writes Off on generated streams, keeps useful real-text learning, and survives broken labels.先用独立证据检验,再提交。Settlement 在生成文本流上几乎追平关闭写入,保留了有用的真实文本学习,并且在标签出错时依然稳健。
Cite引用

BibTeX

@article{luo2026selfgenerated,
  title   = {Self-Generated Feedback Destabilizes Test-Time Training:
             A Causal Decomposition of Long-Horizon Adaptation},
  author  = {Luo, Cheng and Li, Bing and Ghanem, Bernard},
  journal = {arXiv preprint arXiv:2610.05076},
  year    = {2026}
}