AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research reports a zeroth-order training method called Dust that estimates updates by perturbing model activations, rather than calculating gradients through backpropagation. In experiments on transformer language models, the researchers say Dust’s estimates become more aligned with backprop as the population grows; the work does not establish that Dust can replace backprop in large-scale production training.

Q Labs Research has published a report describing Dust, a method for training transformer language models without backpropagation, the standard procedure for calculating how model parameters should change. The team reports that Dust can produce useful training updates in its experiments, but says its closest comparisons with backprop use larger perturbation populations and therefore substantially more compute.

Dust is a zeroth-order optimization method: rather than computing derivatives through a network, it perturbs the model’s activations and uses changes in the loss to estimate an update. The researchers say it applies perturbations independently at each token. This makes tokens act as members of a virtual population that can be evaluated in parallel in a forward pass, without separately creating and running a complete model for each candidate, as conventional weight-space evolution strategies do.

The report says Dust’s gradient estimates become more aligned with backprop as the population grows and remain aligned across the scales tested, with experiments extending up to 1 billion tokens. It also reports that a 243-million-parameter model outperformed a model 120 times smaller at most tested population sizes. These are results described by the authors, not independent confirmation that the method will perform similarly on larger training runs or different tasks.

Q Labs compares Dust with EGGROLL, an evolution-strategy method that perturbs weights. The report estimates Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from 1 million tokens onward. That figure is an extrapolation presented by the researchers; the report’s comparison does not establish an equivalent advantage over backpropagation. The authors also say Dust approximates backprop more closely at large populations and exceeds it in some settings, while acknowledging the greater compute required at those populations.

At a glance
reportWhen: Research report dated October 2026
The developmentQ Labs Research has published experiments on Dust, a method for pretraining transformer language models using activation perturbations instead of backpropagation.

A Different Route to Model Updates

Backpropagation is central to training current deep-learning systems because it efficiently calculates how each parameter contributes to a prediction error. Dust tests whether a search-based approach can train transformers using a different kind of credit assignment: perturbing internal activations and rewarding perturbations associated with lower loss. If methods of this kind become practical, they could broaden the training algorithms available to researchers and reduce reliance on differentiability as a design requirement.

For now, the report establishes a research result, not a practical alternative ready to displace backprop. The central trade-off is compute versus gradient information: larger populations improve Dust’s approximation, but the authors say that takes substantially more compute. The comparison with EGGROLL is promising on the authors’ terms, yet it does not answer whether Dust can match the cost and performance of backprop on production-scale pretraining.

Amazon

transformer language model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Weight Search to Activations

Evolution strategies typically evaluate candidate models created by changing their weights, which can require materializing and running many candidates. Dust instead searches in activation space, perturbing signals inside the network. Its use of independent token-level perturbations is intended to let a single forward pass represent many members of a virtual population.

The report frames this as a test of a broader research question: whether compute-intensive search methods could eventually complement or challenge gradient-based learning. That remains a proposal, not a demonstrated forecast. The work’s stated evidence is a set of experiments with transformer language models and comparisons against backpropagation and EGGROLL.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, in the report’s summary

Amazon

GPU for deep learning training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Scale, Cost and Independent Checks

The report does not establish whether Dust can train models at the scale, duration or data volume used for leading language-model systems. It also does not settle how much compute is needed to match backprop’s training progress across a full pretraining run, or whether the reported advantages persist under different architectures, datasets and hardware configurations.

The authors’ claims about competitiveness, alignment and efficiency are based on their experiments and, for the EGGROLL comparison, extrapolations. The supplied report does not identify independent replications or a production deployment. How much Dust’s approach changes practical training costs, and whether it yields benefits beyond the tested settings, remains unclear.

Amazon

high performance computing for AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Replication and Larger Training Runs

The next evidence to watch for is independent replication and testing on larger, longer training runs. Such work would help establish whether Dust’s token-level virtual population scales beyond the report’s experiments and how its compute requirements compare with backpropagation on equal terms.

The report describes a method and experimental findings; it does not announce a deployment or a scheduled follow-up milestone. Until broader results are available, Dust is best understood as a research test of zeroth-order transformer training rather than a confirmed replacement for the standard backward pass.

Amazon

AI model training accelerators

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Dust?

Dust is a zeroth-order training method that estimates updates by perturbing a model’s activations and measuring changes in the loss, rather than calculating gradients with backpropagation.

Does Dust eliminate backpropagation?

The method is designed to train without a backward pass, and Q Labs reports experiments using it. The findings do not show that Dust has replaced backpropagation in large-scale or production training.

How does Dust differ from evolution strategies?

Conventional weight-space evolution strategies perturb model weights and evaluate candidate models. Dust perturbs activations at each token, which the authors use to create a virtual population evaluated in parallel during a forward pass.

What does the efficiency claim mean?

Q Labs estimates Dust is about 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from 1 million tokens onward. The report calls this an extrapolation; it is not a stated efficiency advantage over backpropagation.

What remains unproven?

Whether Dust can match backpropagation’s practical performance and compute cost in large-scale pretraining remains unclear. Independent replication and tests on larger runs would help answer that.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Memory Squeeze: Why Your RAM Bill Doubled

DRAM prices have surged up to 6x due to a shift in manufacturing focus toward AI hardware, causing shortages and higher costs for consumers.

Is AI causing a repeat of Front end’s Lost Decade?

Analysis of how AI’s impact on programming resembles the 2010s frontend deskilling, raising questions about skill loss, job security, and industry evolution.

Disk Is the Contract: Inside Threlmark’s Local-First Architecture

Discover how Threlmark turns local disk storage into the system’s single source of truth, enabling seamless offline work and powerful AI integration.

Ultra‑Secure Quantum Communication Networks

Prepare to explore the future of ultra-secure quantum communication networks, where your privacy is paramount and secrets are safeguarded like never before.