Multi-Signal Fusion: How Combining Data Sources Can Add Context
Single-source analysis can leave context unresolved. This educational overview explains how FTD, short-volume, options and reference records may be compared without overstating what a combined pattern proves.
By BlueLedger Engineering · 11 min read
The Single-Source Problem
Traditional market surveillance systems were designed around single data sources: one system monitors trade executions, another watches order flow, a third reviews communications. Each system excels at detecting simple, well-defined patterns within its domain - but struggles with complex, multi-dimensional integrity risks that span multiple data types.
This architectural limitation is not a deficiency of any individual system. It reflects the historical reality that surveillance technology evolved incrementally, with each new data source getting its own monitoring silo.
The result is a landscape where individual anomalies are detected reliably, but **correlated patterns across data sources** - often the most meaningful signals - fall through the gaps.
What Multi-Signal Fusion Looks Like
Multi-signal fusion is the practice of ingesting, normalizing and cross-referencing multiple data streams to add context to a review. A combined pattern is not automatically more reliable: source independence, comparability and data quality must be established.
An illustrative review might consider five source classes:
1. Settlement Data (SEC FTD Reports) Twice-monthly SEC disclosures of aggregate failure-to-deliver positions per security. This data provides a settlement observation that may be relevant to a defined review; it is not an integrity finding.
2. Short Volume Data (FINRA) Daily short volume reports from FINRA-regulated trading venues. Short volume as a percentage of total volume provides one context measure, not a conclusion about positioning or pressure.
3. Options Flow Data (Cboe) Options activity data can include volume, open interest, and strike/expiration distributions. An unusual observation has multiple explanations and does not establish an integrity event.
4. Reference Data (ESMA FIRDS, OpenFIGI, GLEIF) Security identification, issuer data and entity-relationship records. Reference data can support cross-jurisdiction analysis when the relevant coverage and mappings are documented.
5. Borrow Cost Indicators Securities lending data that reflects the cost and availability of borrowing shares for short selling. Elevated borrow costs, particularly when combined with high short volume, may indicate constrained supply.
Primary source references: SEC Fails to Deliver Data, https://www.sec.gov/data-research/sec-markets-data/fails-deliver-data; FINRA short-sale volume data, https://www.finra.org/finra-data/browse-catalog/short-sale-volume-data; Cboe market statistics, https://www.cboe.com/us/options/market_statistics/; ESMA FIRDS, https://registers.esma.europa.eu/publication/; OpenFIGI, https://www.openfigi.com/api; and GLEIF LEI data, https://www.gleif.org/en/lei-data/gleif-concatenated-file. Coverage, licensing and definitions must be checked for each analysis.
The Fusion Process
Raw data integration is necessary but not sufficient. The value of multi-signal fusion comes from the analytical layer that sits on top of the integrated data:
**Temporal alignment.** Different data sources operate on different reporting schedules - FTD data is released with a lag, short volume may be daily, and options data may update more frequently. A review must account for those differences before comparing records.
**Baseline calibration.** Each data stream may have its own historical baseline. What constitutes “unusual” for FTDs can differ from short volume, so any threshold and window should be documented.
**Cross-signal comparison.** A reviewer can ask whether multiple comparable records show unusual readings in the same period. A single elevated data point may mean little, and multiple observations still require alternative explanations.
**Confidence scoring.** If a review uses a confidence score, its method should state: - The number of independent signals showing anomalous readings - The statistical significance of each individual deviation - The temporal proximity of the deviations - Historical false positive rates for similar signal combinations
What Fusion Reveals
To illustrate the value of multi-signal analysis, consider a simplified example:
**Single-source view (FTD only):** Security XYZ shows elevated FTDs for three consecutive reporting periods. This is notable but could reflect routine settlement processing delays.
**Single-source view (Short volume only):** Security XYZ shows short volume exceeding 60% of total volume for five consecutive days. This is elevated but within the range of normal market-making activity for some securities.
**Fused view:** Security XYZ shows simultaneous elevation in FTDs, short volume exceeding 60% for five days, unusual put option activity at strikes 30% below current price, and a spike in borrow costs from 1% to 15% annualized - all occurring within the same two-week window. This illustrative convergence provides more context than any individual observation, but does not establish a cause.
The fused example does not prove anything. It illustrates a convergence that could warrant human review, subject to source verification, alternative explanations and the limits of the available data.
Architectural Considerations
Building a multi-signal fusion system requires addressing several engineering challenges:
**Data Quality.** Each data source has its own quirks, errors, and reporting inconsistencies. A robust fusion system must include data quality checks and anomaly detection at the ingestion layer - before signals are generated.
**Latency Management.** Different data sources update at different frequencies. The system must handle mixed-latency inputs without generating false correlations from stale data.
**Scalability.** Cross-referencing multiple streams across many securities requires efficient computation. A production implementation would need to document processing design, freshness and measured performance rather than implying real-time coverage.
**Explainability.** A review should identify which records contributed, what each observation was and how any composite assessment was formed. Availability of these controls is implementation-specific.
The Human Element
Multi-signal fusion is a method for supporting human review, not a replacement for it. Quantitative patterns can suggest questions, but they cannot prove a cause or intent.
An evidence-led interface would present any combined assessment alongside its source records, method and limitations. BlueLedger's current availability and implementation boundaries require separate confirmation.
*BlueLedger provides market integrity monitoring signals and educational content. It is not investment advice and does not allege wrongdoing. Signals indicate anomalies that may warrant review.*