From Market Data to Simulated Fills: Building an End-to-End Trading Pipeline with DolphinDB

DolphinDB
2026-09-29

Computing factors and generating a model signal is only part of the trading workflow. The signal still has to become an order and, eventually, a fill. Between these steps are market-data aggregation, factor joins, model inference, signal generation, order management, and matching.

In high-frequency trading, these stages often process hundreds of millions of records per day. A mismatch in timestamps, schemas, or state at any point can break the pipeline—or, worse, produce backtest results that do not reflect live execution.

Before going live, teams usually build a backtesting system to validate their strategies. In traditional setups, stream processing, model serving, order management, and simulated matching live in separate systems, held together by hand-written glue code that moves data and syncs state. The longer the chain, the more likely you are to run into inconsistent field definitions, misaligned timestamps, out-of-sync state, and simulation logic that drifts away from live trading logic. The simulation results may then say little about how the strategy would actually execute.

To address this, DolphinDB has packaged its real-time computing, matching, inference, and order management components into the SnapTradeStream module. You can download the full module and sample code here.

An End-to-End Streaming Pipeline

SnapTradeStream organizes the key steps from market data to fills into a single streaming pipeline. It covers more than "market data to signal": it extends to signal-to-order and order-to-fill, so you can validate execution as well.

Pipeline Architecture

You only need to write snapshot quotes and tick-by-tick trades into DolphinDB stream tables. The pipeline then handles the rest: market-data processing, factor computation, model inference, order generation, and simulated matching.

Simulated Fill Details

Note: SnapTradeStream depends on the MatchingEngineSimulator plugin and the OrderManagementEngine plugin. Install both from the plugin marketplace before using the module. The simulated matching engine is a paid plugin.

How the Pipeline Works

Each stage's output feeds directly into the next, so you don't move data by hand or write glue code to reconcile formats and timestamps. Here are the stages in order.

Data Cleaning and Time Alignment

Before generating minute-level factors, the module processes the time dimension of both the snapshot and tick-by-tick trade streams. It filters out records from non-trading hours, maps early pre-open quotes to the official market open, and corrects timestamps at minute boundaries. This preserves chronological ordering across batches and prevents window-based calculations from crossing minute boundaries incorrectly.

Building Minute-Level and Derived Factors

Snapshots are aggregated into minute-level market factors: OHLC, volume, turnover, trade count, total bid/ask order quantity, and weighted average bid/ask prices. Trades are aggregated into minute-level trading-flow factors, including active buy and sell amounts, large- and small-order amounts, and trade bias based on the last trade of each minute.

A stream join engine then aligns the two sets of factors by security ID and trade time, and a reactive state engine computes derived factors on the joined stream. The module includes 10 derived factors, covering metrics such as the ratio of turnover to resting orders, large-order turnover against book depth, price deviation, liquidity consumption, and a composite market-pressure factor.

To add your own factors, you define the calculation logic, output field names, and data types in a factor function. The framework creates the underlying engines, state computation, and persistent subscriptions. In most cases, changing a factor means editing that function, not the stream-table schemas or engine definitions.

Model Inference and Signal Generation

As these indicators are generated, they are also persisted and used as inputs to machine learning models, which run inference in real time to produce trading signals.

The module uses two models working together, LightGBM and PyTorch, and both must be configured. After the derived factors pass through the models, the output is a classification (buy, sell, or hold), from which the final strategy signal is generated. The system then turns each strategy signal into an order carrying direction, price, quantity, effective time, and expiry time.

Order Management and Simulated Matching

Signals don't stop at an intermediate results table. They flow on into the order management engine, which maintains each order's lifecycle: creation, status updates, cancellation, partial fills, and rejection records. Each order can be given an effective time and an expiry time. If it is still unfilled when it expires, the system cancels it automatically, so orders never sit on the book indefinitely. This makes the simulation reflect the real-world constraint that orders have a time limit, and makes the results more credible.

The orders then enter simulated matching. The system builds the quote data the matcher needs from real-time snapshots, feeds it into the matching engine together with the strategy's orders, and outputs full fill details: fill price, fill quantity, order status, average fill price, and rejection information. This extends strategy evaluation from "was the prediction accurate?" to execution questions such as "could the order actually fill after entry, and was the price reasonable?" The matching engine also records the amount of unfilled volume ahead of the order and the number of better-priced levels on the book. This makes it possible to trace why an order remained unfilled instead of treating the result as a black box.

Is It Only for Simulation?

Once these four stages are running, a complete simulated fill has gone through the whole chain. That raises a natural question: does the workflow stop at simulation?

As currently implemented, SnapTradeStream provides a validation framework from market data to simulated fills. Its order execution stage relies on the simulated matching engine, so it is not an out-of-the-box live trading system.

That said, its real-time market data processing, factor computation, and model inference can serve as the signal-production layer of a live system. To trade for real, you would connect the order output to a broker API, an OMS, or a trading gateway, replacing the final simulated matching stage with a real execution component.

So SnapTradeStream suits pre-launch work such as historical replay, real-time simulation, fill-quality analysis, and shadow trading. It can also serve as the real-time factor and signal layer of a live trading system.

Performance and Pipeline Simplification

DolphinDB 3.00.4 and later support low-latency computation for streaming engines. Enabling this mode reduces the computation time of derived factors, with the improvement becoming more pronounced as the number of instruments increases.

In testing, when the universe grows from 100 to 1,000 stocks, the reduction in derived-factor computation time becomes increasingly significant, reaching approximately 17× at 1,000 stocks. The benefit is particularly relevant to market-wide, multi-instrument workloads, where both data volume and the number of concurrent computations are large. Actual performance depends on the workload and hardware configuration.

Developer productivity is another consideration. The same pipeline can also be expressed using DolphinDB's Orca stream graph. Instead of manually creating stream tables, defining schemas, configuring engines, and wiring subscriptions together, the pipeline is expressed as a declarative chain.

tradePipeline = g.source(tradeStreamName, 1:0, colNames1, colTypes1)
                .dailyTimeSeriesEngine(...)
                .reactiveStateEngine(...)
                .sink(factorDbname+"/"+tradeMinStreamTbname)

This approach makes it easier to add, remove, or reorder processing stages without rebuilding the surrounding stream infrastructure. For teams iterating frequently on factors or strategy logic, it reduces repetitive pipeline configuration and keeps more of the implementation focused on the strategy itself.

Conclusion

A trading strategy is not fully validated when its signal performs well on historical data. The signal still has to become an order, interact with the order book, and produce a realistic fill.

By connecting market-data processing, factor computation, model inference, order management, and simulated matching in a single pipeline, SnapTradeStream extends strategy validation from signal quality to execution quality. Instead of evaluating only whether a prediction was correct, teams can also examine whether the resulting order could have been filled, at what price, and under what liquidity conditions.

This end-to-end workflow helps narrow the gap between backtesting and live execution. The same real-time data, factor, and inference pipeline can also be reused as the signal-production layer of a live trading system, with the simulated matching stage replaced by a production execution gateway.