Play with swarm scaling动手玩一玩 swarm scaling
Before the mechanics, try the button below. A task forms a directed acyclic graph of steps, a DAG, because each step builds on earlier steps and can start only once they are finished. Each press draws a small random task DAG, grown the same way as the DAGs in this post, and runs it with one agent and then with standard swarm@4. In a standard swarm, a coordinator starts a new agent on a ready step whenever an agent runs out of ready steps. The left chart rates each run by coverage, the share of steps finished. The right chart rates it by the best score, the highest value among those steps.
在讲机制之前,先点一下下面的按钮。任务可以看作由步骤构成的有向无环图,即 DAG:每一步都建立在之前的步骤之上,要等它们完成后才能开始。每点一次,都会按本文的方法随机生成一张小的任务 DAG,先由单个智能体完成,再由标准蜂群@4 完成。在标准蜂群中,每当有智能体做完了自己列表里的所有就绪步骤,调度器就派一个新的智能体去做某个就绪的步骤。左图按覆盖度评估每次运行,即已完成步骤的比例;右图按最优分数评估,即已完成步骤中的最高分值。
Task DAGs and one agent任务 DAG 与单个智能体
An autonomous agent explores an unknown environment one step at a time, and each finished step opens new ones. Over a long run its progress follows a predictable curve1. We model the environment as a DAG of such steps and first let one agent work through it.
自主智能体一步一步地探索未知环境,每完成一个步骤都会打开新的步骤。在长时间运行中,它的进度沿一条可预测的曲线推进1。我们把环境建模为由这些步骤组成的 DAG,先让单个智能体在上面探索。
The task DAG任务 DAG
Each step is a node, and an edge runs from a prerequisite to the step it unlocks. A step becomes ready only when all of its prerequisites are finished. Deeper steps generally take longer. The DAGs in this post grow by trial and error, as in a search for a pretraining recipe. Each new step is a variant of a parent in the layer above, and the more a parent improved on what it built on, the more often it is drawn. Some steps also merge extra parents, the way a recipe combines earlier changes. A step's value is its best parent's value plus a random change. The animation below grows a small example.
每个步骤是一个节点,边从前置步骤指向它解锁的步骤。一个步骤只有在所有前置步骤都完成后才会就绪。越深的步骤通常越耗时。本文的 DAG 由反复试错生长而成,就像搜索预训练 recipe 的过程:每个新步骤都是上一层某个父步骤的变体,父步骤相对其前身改进越大,被选中的机会就越多。有些步骤还会合并额外的父步骤,就像一个 recipe 把之前的几项改动组合起来。步骤的分值等于其最好的父步骤的分值加上一个随机变化。下面的动画生长出一个小例子。
How the generator works生成器如何工作
- A task DAG has about 723 steps in 17 layers, the size of the language modeling DAG. Layer 0 holds only the root, the baseline recipe with value 0. The sizes $m_d$ of the other layers follow a bell over depth that peaks 28% of the way down: layer $d$ gets the rise of a logistic curve of steepness $6.4/(D-1)$ centered at the peak across that layer, times a noise factor $e^{0.35 z}$, where $z$ is drawn from a standard normal each time it appears.一张任务 DAG 约有 723 个步骤、17 层,与语言建模 DAG 的规模相同。第 0 层只有根节点,即分值为 0 的基线 recipe。其余各层的步骤数 $m_d$ 沿深度呈钟形,峰值位于 28% 深度处:第 $d$ 层分到一条以峰值为中心、陡度为 $6.4/(D-1)$ 的 logistic 曲线在该层上的增量,再乘以噪声因子 $e^{0.35 z}$。本文中 $z$ 每次出现都从标准正态分布中重新抽取。
- Each figure runs 16 tasks, and each task varies around these values: it has $723\,e^{0.4 z}$ steps, $17\,e^{0.2 z}$ layers, its peak at $0.28 + 0.07 z$ of the depth and a merge probability $p = 0.46\,e^{0.3 z}$, each with a $z$ of its own. Its clock is also stretched by $e^{\delta_b}$, with $\delta_b$ uniform between $-0.3$ and $0.3$, so that tasks of one size still differ in speed.每张图使用 16 个任务,每个任务在上述取值附近变化:步骤数为 $723\,e^{0.4 z}$,层数为 $17\,e^{0.2 z}$,峰值位于 $0.28 + 0.07 z$ 的深度处,合并概率 $p = 0.46\,e^{0.3 z}$,每项各用一个独立的 $z$。任务的时钟还要乘以 $e^{\delta_b}$,$\delta_b$ 在 $-0.3$ 到 $0.3$ 之间均匀分布,因此规模相同的任务速度也不完全一样。
- Layer by layer, each new step draws its main parent from the layer above with probability proportional to $e^{\beta \delta}$, where $\beta = 1.4$ and $\delta$ is that parent's own improvement over what it built on. A step that improved on its parent thus collects more of the variants, whatever its lineage.DAG 逐层生成:每个新步骤从上一层抽取主父步骤,概率与 $e^{\beta \delta}$ 成正比,其中 $\beta = 1.4$,$\delta$ 是该父步骤相对其前身的改进。因此,一个步骤只要比它的前身有所改进,就会得到更多变体,与它的来历无关。
- With probability $p$, a step also merges $K$ extra parents, with $P(K) \propto K^{-2.3}$ for $K$ from 1 to 16.每个步骤还以概率 $p$ 合并 $K$ 个额外父步骤,$P(K) \propto K^{-2.3}$,$K$ 取 1 到 16。
- Each extra parent is a sibling of the main parent with probability 0.2. Otherwise it comes from the layer $b$ levels up, with $P(b) \propto 0.8^{b}$ as far as the root, and is drawn there with the same weights. Two parents of one step never lie on the same chain: a candidate that is an ancestor of a chosen parent is skipped, and a chosen parent that is an ancestor of a new one is dropped.每个额外父步骤以 0.2 的概率取主父步骤的兄弟步骤;否则取自向上第 $b$ 层,$P(b) \propto 0.8^{b}$,最远到根节点,并在该层按同样的权重抽取。同一步骤的两个父步骤不会位于同一条链上:若候选步骤是某个已选父步骤的祖先,就跳过它;若某个已选父步骤是新选步骤的祖先,就把它去掉。
- A new step's value is its best parent's value plus a change drawn from $\mathcal{N}(\mu_d, 1)$, where the mean $\mu_d$ falls linearly with depth from $+1$ at the root to $-0.5$ in the deepest layer. Near the root about five variants in six beat their parent, and in the deepest layer about one in three, so gains come early and the best score moves only through the few deep variants that still improve.新步骤的分值等于其最好的父步骤的分值加上一个服从 $\mathcal{N}(\mu_d, 1)$ 的变化,均值 $\mu_d$ 随深度线性下降,从根节点处的 $+1$ 降到最深一层的 $-0.5$。靠近根节点时约六分之五的变体优于父步骤,最深一层只有约三分之一,因此收益集中在前期,后期最优分数只靠少数仍有改进的深层变体提升。
- Each DAG also draws a fertility $q$ from $\mathcal{N}(0, 0.2^2)$ and adds it to the mean of every change in it. Independent attempts at one task thus take directions of different promise, and a DAG that starts well tends to end well. A shift shared by every step of a layer leaves the parent weights as they were, so the fitted shape does not change.每张 DAG 还抽取一个肥沃度 $q \sim \mathcal{N}(0, 0.2^2)$,加到其中每一次变化的均值上。于是同一任务的几次独立尝试会走上前景不同的方向,开局好的 DAG 往往结局也好。同一层所有步骤平移同样的量,不改变父步骤的抽取权重,所以拟合出的结构不变。
- A step's cost depends on its layer alone and rises linearly from 1 at the root to 10 in the deepest layer, $c_d = 1 + 9d/(D-1)$. An agent's time on a step is that cost times $e^{0.2 z}$.步骤的代价只取决于所在层,从根节点的 1 线性增加到最深一层的 10,即 $c_d = 1 + 9d/(D-1)$。智能体完成一个步骤的实际耗时是该代价乘以 $e^{0.2 z}$。
- Every DAG of a task, its own and those that independent attempts explore, is scored on one scale. We grow 64 DAGs like the task's and let the 90th percentile of their best recipes' values be worth 1, and a higher value counts as 1. On this scale the best recipe of a typical DAG is worth about 0.83, and about one DAG in six reaches 0.95.同一任务的所有 DAG,包括它自己的和各次独立尝试探索的,都用同一个尺度打分。我们按该任务的方式生成 64 张 DAG,把它们最优 recipe 分值的第 90 百分位记为 1,更高的值也记为 1。在这个尺度上,一张典型 DAG 的最优 recipe 约为 0.83,大约六张里有一张能到 0.95。
- The exponent 2.3, the sibling share 0.2, the decay 0.8 and $\beta = 1.4$ are fitted to two real exploration DAGs. They come from some internal recursive self-improvement experiments, one searching for a language modeling pretraining recipe and one for a text-to-image diffusion recipe. The fitted generator matches their share of steps with each number of parents, how often merges join siblings and how unevenly children spread across parents. With the means $\mu_d$ above, the best recipe of a generated DAG sits at 94% of its depth at the language modeling DAG's size and at 88% at the diffusion DAG's size, against 94% and 92% in the real DAGs.指数 2.3、兄弟比例 0.2、衰减率 0.8 和 $\beta = 1.4$ 都拟合自两张真实的探索 DAG,它们来自我们内部的一些递归自我改进实验,一个搜索语言建模的预训练 recipe,一个搜索文生图 diffusion 的 recipe。拟合后的生成器在三方面与这两张 DAG 一致:不同父步骤数的步骤占比、合并发生在兄弟步骤之间的频率,以及子步骤在父步骤间分布的不均匀程度。取上面的均值 $\mu_d$ 时,生成的 DAG 中最优 recipe 位于 94%(语言建模 DAG 的规模)和 88%(diffusion DAG 的规模)深度处,真实 DAG 中分别为 94% 和 92%。
- An extra parent from $b \ge 2$ layers up adds a cross-layer edge, for example from layer 1 into layer 4. A real DAG has no layers of its own, so we place each of its steps one layer below its deepest parent, as the generator does. Measured this way, cross-layer edges make up 26% of the edges in the language modeling DAG, 16% in the diffusion DAG and 18% in the generated DAGs. On this count the generator comes close to the diffusion DAG, and the language modeling DAG has more cross-layer edges than either.来自向上 $b \ge 2$ 层的额外父步骤会形成一条跨层边,例如从第 1 层连到第 4 层。真实 DAG 本身没有分层,我们像生成器一样,把每个步骤放在其最深父步骤的下一层。按这种方式统计,跨层边在语言建模 DAG 中占全部边的 26%,在 diffusion DAG 中占 16%,在生成的 DAG 中占 18%。就这一指标而言,生成器接近 diffusion DAG,语言建模 DAG 的跨层边则比两者都多。
Two scores for two kinds of exploration两类探索对应的两种分数
A finished step can be scored in two ways, one for each kind of exploration. In discovery, an agent learning a codebase or a protocol wants every fact. Each step is one fact worth one point, so the score is coverage, the share of steps finished. In optimization, an agent searching for a pretraining recipe keeps only its best result. Its best score is the highest value among finished steps, on one scale for every DAG of the task, where 1 is a best recipe that one DAG in ten reaches. The other nine fall short of it, so on one DAG the average best score levels off below 1, even after every step is finished. Coverage counts every branch, while the best score waits on the one chain that leads to the best recipe, deep in the DAG. Figures 1 to 4 plot both side by side, and Figure 5 chooses the best swarm by the best score. Time is measured in units of T₁, how long the slowest single-agent run takes to finish every step.
已完成的步骤有两种计分方式,分别对应两类探索。在发现型探索中,比如熟悉一个代码库或一套协议,智能体需要掌握每一条事实:每个步骤是一条事实,记一分,因此分数就是覆盖度,即已完成步骤的比例。在优化型探索中,比如搜索预训练 recipe,智能体只保留最好的结果:最优分数是已完成步骤中的最高分值,同一任务的所有 DAG 共用一个尺度,1 对应十张 DAG 里有一张能达到的最优 recipe。其余九张达不到这个水平,所以在单张 DAG 上,即使做完全部步骤,平均最优分数也停在 1 以下。覆盖度计入每一条分支,最优分数则取决于通往最好 recipe 的那一条链,它位于 DAG 深处。图 1 到图 4 同时画出这两种分数,图 5 按最优分数选出最好的蜂群。时间以 T₁ 为单位,即最慢的一次单智能体运行完成全部步骤所需的时间。
Single-agent baseline单智能体基线
An agent works through the DAG one step at a time. It keeps a list of the ready steps one of whose parents it finished and always takes the deepest of them, the one furthest along. After a step that makes new steps ready, that is usually one of them, so the agent keeps going down its branch. After a dead end, a step that makes no new step ready, it goes back to the deepest step left on its list. Every agent in this post follows this rule, alone or in a swarm. The animation below follows one agent on a small DAG.
智能体每次做一个步骤。它维护一个列表,记录自己完成过其某个父步骤的就绪步骤,每次取其中最深、也就是走得最远的那个。如果刚做完的步骤让新的步骤就绪,最深的通常就是其中之一,于是它沿这条分支继续往下;如果走进了死胡同,即没有新步骤就绪,它就回到列表里剩下的最深的步骤。本文所有智能体,无论单独工作还是在蜂群中,都遵循这条规则。下面的动画展示一个智能体走完一张小 DAG 的过程。
The step rule in detail选步规则的细节
- A step is ready once all of its parents are finished, so a step that merges extra parents waits for every one of them.一个步骤只有在所有父步骤都完成后才会就绪,因此合并了额外父步骤的步骤要等它们全部完成。
- A ready step is on an agent's list once the agent has finished any of its parents, so a merge step can be on several lists and goes to whichever agent takes it first. The agent takes the deepest step on its list, and among the steps of one layer it takes one at random.一个就绪步骤只要有一个父步骤是某个智能体完成的,就会进入这个智能体的列表,因此合并步骤可能同时出现在几个列表里,谁先取到归谁。智能体每次取列表中最深的步骤,同一层的步骤之间随机选一个。
- About half the steps of a DAG have no children, and an agent cannot tell such a step before it tries it.一张 DAG 中约一半的步骤没有子步骤,智能体在尝试之前无法识别这样的步骤。
- In both swarms, an agent that empties its list stops, and a freed slot goes to a random ready step. The running agent that made that step ready, by finishing its last parent, holds it. In a standard swarm the coordinator starts a new agent on it. In a recursive swarm that agent forks a sub-agent for it instead, unless its layer is at the cap, and an agent whose finished step opens several steps also keeps one of them and forks sub-agents for the others while the budget has room.在两种蜂群中,列表清空的智能体都会停止,空出的名额随机分给一个就绪步骤。完成了这个步骤最后一个父步骤、让它就绪的那个在岗智能体持有它。标准蜂群由调度器派一个新智能体去做它;递归蜂群则由持有它的智能体派生一个子智能体去做,除非该智能体已处于层数上限。此外,在递归蜂群中,如果一个智能体做完的步骤打开了多个新步骤,它会保留其中一个,并在预算允许时为其余步骤各派生一个子智能体。
- pass@k runs k agents that each follow this rule on their own.pass@k 让 k 个智能体各自独立地遵循这条规则。
How swarms divide work蜂群如何分工
To speed up exploration, several agents can work on ready steps in parallel. We compare two ways to organize a swarm, both run by a coordinator: a standard swarm, whose agents all report to the coordinator, and a recursive swarm, whose agents also fork sub-agents of their own. A swarm's speedup g_x is how many times sooner it reaches level x than one agent. Its agent budget, set by the sliders, is the most agents it may run at once. A swarm with a budget of n agents is written standard swarm@n or recursive swarm@n. Every swarm in this post pays the costs set out next.
为了加快探索,可以让多个智能体并行处理就绪的步骤。我们比较两种组织蜂群的方式,两者都由一个调度器管理:标准蜂群中所有智能体都向调度器汇报;递归蜂群中的智能体还会派生自己的子智能体。蜂群的加速比 g_x 是单个智能体与蜂群到达水平 x 所用时间之比。智能体预算由滑块设定,是蜂群同时运行的智能体数上限;预算为 n 的蜂群记作标准蜂群@n 或递归蜂群@n。本文中的每个蜂群都承担下一节设定的开销。
Real-world constraints现实中的制约
Working as a swarm introduces two sources of overhead. Scheduling overhead is the one-time cost of handing over context: 5% of a median step whenever an agent starts or builds on another agent's work. Dispatching a step takes the coordinator no time. Communication cost is the continuous cost of keeping agents in sync, and each group pays it. A group is a superior and the agents directly under it, and the coordinator, which does no steps of its own, is not counted. While two agents of a group are active, each of their steps takes 15% longer, and every further active member adds 1% to every step in the group. In a standard swarm every agent reports to the coordinator, so all its agents form one group. In a recursive swarm an agent with sub-agents of its own sits in two groups, with its peers and with its sub-agents, and pays for both.
以蜂群方式工作会带来两类开销。调度开销是交接上下文的一次性代价:每当一个智能体启动,或在其他智能体的成果上继续工作时,都要多花一个中位步骤耗时的 5%。调度器派发步骤不花时间。沟通成本是保持智能体之间同步的持续代价,按小组承担。一个小组由一个上级和直接归它管的智能体组成;调度器自己不做步骤,不计入人数。小组中只要有两个智能体在岗,它们的每一步都多花 15% 的时间;此后每多一个在岗的组员,组内所有人的每一步再多花 1%。标准蜂群中每个智能体都向调度器汇报,所以全体智能体是一个小组。递归蜂群中,有子智能体的智能体同时属于两个小组,一个是它和同级的智能体,一个是它和自己的子智能体,两份开销都要承担。
The animation below runs both swarms with four agents. With all four at work, every step in the standard swarm takes 17% longer: 15% for the group and 1% for each of the two members beyond the second. In the recursive swarm the coordinator starts two agents here, and each of them forks one sub-agent. A dispatched agent sits in two groups of two, one with the other dispatched agent and one with its sub-agent, and pays 30%, and each sub-agent pays 15%.
下面的动画让两种蜂群各运行四个智能体。四个智能体都在工作时,标准蜂群中的每一步要多花 17%:小组本身占 15%,第二个之后的两个组员各占 1%。在这里的递归蜂群中,调度器启动两个智能体,它们各派生一个子智能体。被派出的智能体同时属于两个两人小组,一个与另一个被派出的智能体组成,一个与自己的子智能体组成,要多花 30%;每个子智能体多花 15%。
Standard swarms标准蜂群
A standard swarm caps the number of concurrent agents. Each agent works until its list of ready steps is empty, then stops and frees its slot. The coordinator immediately starts another agent on a random ready step.
标准蜂群限定同时运行的智能体数量。每个智能体一直做到自己的就绪列表清空,然后停止并释放名额,调度器随即派一个新智能体去做随机选出的就绪步骤。
Both curves shift earlier as the swarm grows. With 64 agents a standard swarm reaches half coverage 33 times sooner than one agent and a best score of 0.8 16 times sooner. From 32 agents on, its last steps come no sooner: it finishes every step at about 0.10 T₁ with 32 agents and with 64, since the longest chain of steps must be worked one step after another.
随着蜂群扩大,两条曲线都整体提前。64 个智能体的标准蜂群到达一半覆盖度比单个智能体快 33 倍,最优分数到达 0.8 快 16 倍。从 32 个智能体起,最后几个步骤不再提前:32 个和 64 个智能体都在约 0.10 T₁ 做完全部步骤,因为最长的那条步骤链只能一步接一步地完成。
Recursive swarms递归蜂群
A recursive swarm keeps the standard swarm's coordinator, but its agents hand out their own extra steps. Whenever an agent's finished step opens more than one new step, the agent keeps one and forks a sub-agent for each of the others, up to the swarm's agent budget, and when a slot frees later, the agent holding the ready step it goes to forks a sub-agent for that step. The coordinator starts the first agent and dispatches only the steps held at the deepest allowed layer. Each sub-agent works through its own branch to the end and may fork in turn.
递归蜂群保留标准蜂群的调度器,但多出来的步骤由智能体自己分派。每当一个智能体做完的步骤打开了不止一个新步骤,它就保留一个,并为其余每个步骤派生一个子智能体,总数不超过蜂群的预算;之后有名额空出时,由持有所分到的就绪步骤的智能体为它派生子智能体。调度器只启动第一个智能体,以及派发处在最深允许层的智能体所持有的步骤。每个子智能体一直做到自己的列表清空,途中也可以继续派生。
How a recursive swarm forks递归蜂群如何派生
- Each agent keeps its own list of ready steps: the step it was forked for and the ready steps whose parents it finished. It works through the list one step at a time, deepest first.每个智能体有自己的就绪步骤列表,包括它被派生时分到的步骤,以及它完成过其某个父步骤的就绪步骤。它每次做一步,先做最深的。
- When a finished step unlocks several new steps, the agent keeps one and forks a sub-agent for each of the others while the budget has room. The steps it cannot hand out join its list.当做完的步骤解锁了多个新步骤时,智能体保留一个,并在预算允许时为其余每个步骤派生一个子智能体;分派不出去的步骤加入它自己的列表。
- A sub-agent that empties its list exits and frees its slot. The slot goes to a random ready step, as in a standard swarm, and the agent holding that step forks a sub-agent for it, unless it sits at the deepest allowed layer, in which case the coordinator dispatches it.列表清空的子智能体会退出并释放名额。和标准蜂群一样,这个名额随机分给一个就绪步骤,由持有该步骤的智能体为它派生子智能体;如果持有者已处于最深允许层,则由调度器派发。
- An agent the coordinator starts reports to the coordinator, as in a standard swarm. A forked sub-agent is the direct subordinate of the agent that forked it and pays the usual costs: the scheduling overhead on its first step, and on every step the communication cost of its group, the agent that forked it and that agent's other sub-agents.由调度器启动的智能体向调度器汇报,这一点与标准蜂群相同。派生出的子智能体是派生它的智能体的直接下属,照常承担开销:第一步的调度开销,以及每一步所在小组的沟通成本,小组由派生它的智能体和该智能体的其他子智能体组成。
- A layer cap keeps the agents of the deepest allowed layer from forking. With one layer only the coordinator starts agents, and the recursive swarm is a standard swarm.层数上限禁止最深允许层的智能体继续派生。只有一层时,只有调度器能启动智能体,递归蜂群就退化为标准蜂群。
- Since every freed slot is filled at once, by a fork or by the coordinator, no budget sits idle while ready steps wait.每个空出的名额都会立即由派生或调度器补上,因此只要有就绪步骤在等待,就不会有预算闲置。
Swarm arena蜂群擂台
This section compares ways to organize k agents. pass@k runs k agents on their own, each on a DAG of its own, and keeps the best score among them. standard swarm@k and recursive swarm@k put all k agents in one swarm, and pass@m standard swarm@n or pass@m recursive swarm@n splits them into m independent swarms of n = k/m agents. Its recursive swarms have three layers, the setting Figure 4 starts at. The agents of one swarm share a DAG. Independent attempts at a real task rarely meet and take different routes, so each independent part explores its own DAG drawn from the task's generator, which models the diversity of independent exploration. A system's best score is the best of its parts'. Figure 5 races the systems two at a time on their average best scores over 4,096 runs. Its second slider sets the task width, the average number of steps per layer: each task keeps its steps and is regrown in about steps / w layers. The first two are drawn at random. The one that reaches 0.95 sooner wins, or 0.8 sooner when neither reaches 0.95, and stays on; the loser leaves, and a system that has not raced yet takes its place. When every system has raced, the draw starts over.
本节比较组织 k 个智能体的几种方式。pass@k 让 k 个智能体各自在自己的 DAG 上独立工作,取其中的最高分。标准蜂群@k 和递归蜂群@k 把 k 个智能体组成一个蜂群;pass@m 标准蜂群@n 和 pass@m 递归蜂群@n 则把它们分成 m 个相互独立、各有 n = k/m 个智能体的蜂群。这里的递归蜂群用三层,与图 4 的默认设置相同。同一个蜂群中的智能体共享一张 DAG。在真实任务中,相互独立的尝试很少相遇,走的路线也各不相同,所以每个独立部分各自探索一张由该任务的生成器得到的 DAG,以此模拟独立探索的多样性;系统的最优分数取各部分中的最高值。图 5 让各种系统两两比赛,比的是 4,096 次运行的平均最优分数。第二个滑块设定任务宽度,即平均每层的步骤数:每个任务保留原来的步骤数,按约“步骤数 / w”层重新生成。第一场的两个系统随机抽取。先到 0.95 的一方获胜,两方都到不了 0.95 时比谁先到 0.8;胜者留下,输的一方下场,由一个还没上过场的系统补上。所有系统都上过场后,重新抽签。
Adaptive swarms自适应蜂群
Every structure in the Swarm arena keeps its shape for the whole run. An adaptive swarm changes shape as it goes: it starts with every agent on its own, then moves agents from the independent parts that lag to the parts that lead, which grow into recursive swarms. It decides from the best score each part has found so far, since some DAGs of a task are richer than others and show it early. Other adaptive algorithms remain to be explored.
蜂群擂台里的每种结构在整个运行中都保持不变。自适应蜂群则边跑边改变结构:开始时每个智能体各自独立探索,之后把智能体从落后的独立部分转到领先的部分,让领先的部分长成递归蜂群。它依据每个部分目前找到的最优分数做决定,因为同一任务的某些 DAG 比其他的更肥沃,而且很早就显现出来。其他自适应算法还有待探索。
How an adaptive swarm works自适应蜂群如何运作
- It starts as pass@k: each of the k agents explores a DAG of its own, as in the Swarm arena.开始时是 pass@k:k 个智能体各自探索一张 DAG,和蜂群擂台中一样。
- It watches one signal, the best score each independent part has found so far, on the task's one scale. A real deployment can rank its parts by this signal, but it does not know where 1 lies. The signal is worth watching because DAGs differ in fertility: a part's best score at 0.08 T₁ already ranks the ceilings of the DAGs with a rank correlation of about 0.4.它只看一个信号:每个独立部分目前找到的最优分数,按该任务统一的尺度计算。真实部署可以按这个信号给各部分排名,但并不知道 1 在哪里。这个信号值得看,是因为各张 DAG 的肥沃度不同:在 0.08 T₁ 时,一个部分的最优分数与其 DAG 上限的秩相关已约为 0.4。
- It halves the parts still running $\log_2 (k/4)$ times, so that four parts are left: each time the leading part's best score first reaches the next of these levels, spread evenly from 0.25 to 0.6, it ranks the parts by their best scores and keeps the upper half. With 16 agents the levels are 0.25 and 0.6. They were picked by a small search on 16 and 64 agents. Four parts are kept because one or two often sit on DAGs whose best recipe falls short of 0.95, and levels set higher merge too late to catch up.它把仍在运行的部分减半 $\log_2 (k/4)$ 次,最后剩下四个部分:领先部分的最优分数每次首次到达下一个水平时,就按最优分数排名,保留前一半。这些水平在 0.25 到 0.6 之间均匀分布,16 个智能体时为 0.25 和 0.6。它们是在 16 和 64 个智能体上小范围搜索得到的。之所以保留四个部分,是因为只剩一两个部分时,它们所在的 DAG 往往连最优 recipe 都到不了 0.95;水平定得更高时,合并又太晚,追不回来。
- The other half stop. What they found still counts toward the swarm's best score, and their agents join the kept parts, so every kept part ends up with the same number of agents.另一半停下。它们找到的结果仍计入整个蜂群的最优分数,它们的智能体加入被保留的部分,每个被保留的部分分到同样多的智能体。
- Each kept part restarts as a recursive swarm of three layers with all its agents, old and new, from the steps it has already finished. A step in progress at a halving goes back among the ready steps, and the time already spent on it is lost.每个被保留的部分带着新旧所有智能体,从已完成的步骤出发,重新作为三层的递归蜂群运行。减半时正在进行的步骤回到就绪步骤中,已经花在它上面的时间就浪费了。
- With 16 agents the structure goes from 16×1 to 8×2 and then 4×4, four recursive swarms of four agents. With 64 agents it goes from 64×1 to 4×16 in four halvings, and with 4 agents it stays 4×1, which is pass@4.16 个智能体时,结构依次为 16×1、8×2、4×4,最后是四个各含 4 个智能体的递归蜂群;64 个智能体时经过四次减半从 64×1 变到 4×16;4 个智能体时保持 4×1,也就是 pass@4。
- Its recursive swarms pay the costs set out in Real-world constraints, like every swarm in this post, and an agent on its own pays none.它的递归蜂群和本文中所有蜂群一样,承担“现实中的制约”一节设定的开销;独自工作的智能体没有开销。
Takeaways核心结论
Under suitable conditions, we can use simulation to roughly estimate the best structure for an agent swarm and its scaling law.
在适当的条件下,我们可以通过模拟大致推算最佳的 agent swarm 结构和 scaling law。
Acknowledgements致谢
We thank Ang Cao, Tony Chen, Mingyang Deng, Wentao Guo, Hanchen Li, Kaiyuan Liu, Qiuyang Mang, Kaiyue Wen and Ziqian Zhong for helpful discussions and feedback on this post (names in alphabetical order). We also thank Karthik Narasimhan for his guidance and support.
感谢 Ang Cao、Tony Chen、Mingyang Deng、Wentao Guo、Hanchen Li、Kaiyuan Liu、Qiuyang Mang、Kaiyue Wen 和 Ziqian Zhong 对本文的有益讨论与反馈(按姓氏字母排序)。同时感谢 Karthik Narasimhan 教授的指导与支持。
References参考文献
Citation
@misc{chai2026predictableswarmscaling,
title = {Predictable Swarm Scaling},
author = {Chai, Wenhao},
year = {2026},
howpublished = {Blog post},
url = {https://wenhaochai.com/blogs/predictable-swarm-scaling.html}
}