NGC 3079 XMM large-scale thermal wind decomposition
algorithm tournament & production pipeline progress

X-ray spectral–morphological component separation on XMM-Newton EPIC data
New methods based on photon energy, PSF morphology, spatial scale, and component-specific vignetting

35EPIC event files
12ObsIDs
23response triplets
5candidate methods
3baseline methods
59FXT ObsIDs

Science goal

Develop, compare, and validate new, generalizable X-ray spectral–morphological component separation algorithms. These algorithms jointly exploit photon energy, PSF morphology, spatial scale, and component-specific exposure/vignetting operators to recover an auditable, physically interpretable 2-D morphology map of the NGC 3079 large-scale thermal wind (thermal wind) from XMM-Newton EPIC raw counts-space data.

Target physical quantities include: wind axis direction (PA), opening angle (opening angle), radial extent (radial extent), north–south asymmetry (N/S asymmetry), surface brightness profile (surface brightness profile), and component flux (component flux), while separating PSF-scale point sources and component-specific backgrounds.

NGC 3079 serves as the primary real-data demonstration and benchmarking target — not a hardcoded cone template. Ultimately the validated morphological forward-folding (forward-fold) will be applied to real FXT data through Einstein Probe FXT responses and backgrounds, producing a calibrated detection/constraint or a strict upper limit.

Scientific methodology

The core innovation is spectral–morphological separation (spectral-morphological separation): rather than relying on a single information channel, four physical dimensions are jointly exploited to separate the diffuse wind component, point sources, and backgrounds.

① Photon energy (photon energy)

Exploits the spectral differences between components (thermal wind, power-law point sources, QPB background) to identify components via multi-band response folding.

② PSF morphology (PSF morphology)

Point sources show PSF-scale morphology while the diffuse wind shows large-scale structure. PSF convolution operators separate point sources from extended emission.

③ Spatial scale (spatial scale)

The wind component extends on kiloparsec scales while point sources concentrate on arcsecond scales. Multi-scale decomposition (wavelet/multiscale) provides the scale separation.

④ Component-specific vignetting (component-specific vignetting)

Effective areas of different energy bands and cameras (MOS1/MOS2/PN) vary differently with off-axis angle. Correctly modeling component-specific vignetting is the precondition for unbiased flux recovery.

Two inference modes

Discovery mode: no priors on the NGC 3079 wind's PA, opening angle, or cone masks are injected. The method must discover the diffuse structure from energy + PSF + scale information alone. The known geometry is used only for blind scoring after inference.

Characterization mode: only after Discovery mode is validated may the recovered or known axis/cone families be used to sharpen PA, extent, asymmetry, and flux constraints.

Algorithm tournament results

Eight methods (5 candidates + 3 baselines) were benchmarked uniformly on 72 synthetic scenarios (36 bipolar cones + 27 symmetric halos + 9 no-wind). Each method was evaluated on 5 metrics and judged against pre-declared kill criteria.

Method Type Leakage Flux Bias PA Bias (°) Extent Bias FP Rate Verdict
ILC (candidate) candidate 0.000 -6.977 0.0 1.000 0.000 INVALIDATED
Poisson (candidate) candidate 0.000 -1.000 0.0 1.000 0.000 INVALIDATED
Dictionary (candidate) candidate 0.934 4.270 70.14 1.302 0.000 INVALIDATED
Semi-blind (candidate) candidate 0.827 1.022 66.90 0.981 0.000 INVALIDATED
Sparse+Smooth (candidate) candidate 0.848 2.515 60.84 1.135 0.000 INVALIDATED
Broad-band (baseline) baseline 0.938 13.777 0.0 0.000 0.000 INVALIDATED
Point-source mask (baseline) baseline 0.933 12.677 0.0 0.000 0.000 INVALIDATED
Hardness map (baseline) baseline 0.975 3.809 0.0 0.000 0.000 INVALIDATED

Leakage metric visualization

Lower is better. Target ≤ 10%.

ILC
0.000
Poisson
0.000
Semi-blind
0.827
Sparse+Sm
0.848
Dictionary
0.934
Broad-band
0.938
PS mask
0.933
Hardness
0.975

Key science findings

ILC and Poisson achieve zero leakage

Both constrained ILC and Poisson profile-likelihood achieved zero leakage (0.000) across all 72 scenarios, far better than the 10% acceptance criterion. This is the critical first step of component separation — point-source flux is successfully prevented from leaking into the diffuse wind component. However, both were INVALIDATED by flux bias (flux bias: ILC -6.977, Poisson -1.000), showing that while leakage-free they over-suppress the wind component's true flux.

Dictionary / Semi-blind / Sparse+Smooth show high leakage and PA bias

spectral-template × wavelet/PSF dictionary (leakage 0.934, PA bias 70.1°), semi-blind physical-template factorization (leakage 0.827, PA bias 66.9°), and sparse-point + smooth/multiscale unmixing (leakage 0.848, PA bias 60.8°) recover part of the wind flux, but with severe point-source leakage and position-angle biases far beyond the 30° kill threshold. In the pilot implementation these methods cannot reliably distinguish wind structure from point-source scattering.

All methods INVALIDATED on the synthetic benchmark — the expected pilot-stage outcome

All 8/8 methods are INVALIDATED. This is the expected and scientifically honest outcome: the pilot implementation used simplified templates and generic physics grids, not target-specific spectral bases. The kill criteria were pre-declared before the tournament began and not modified after the verdicts. The result gives the P4 algorithm improvements and hybrid schemes clear quantitative directions: ILC/Poisson must fix their flux bias, and the Dictionary family must solve its leakage and PA bias.

The production pipeline (P3) is connected to real EPIC data

The P3 production line implements: real EPIC data ingestion (35 science-filtered event files, 12 ObsIDs, M1/M2/PN three cameras) → response folding (23 complete ARF+RMF+PI triplets, OGIP-standard validated) → exposure correction (component-specific vignetting operators) → uncertainty propagation (uncertainty maps module). This means the algorithm tournament's findings can directly drive the real-data decomposition.

Production pipeline architecture

The complete data flow from raw XMM-Newton EPIC counts to science measurements:

P0 Input inventory & provenance │ 35 EPIC event files · 12 ObsIDs · 37 event primary keys │ 23 ARF+RMF+PI triplets (OGIP validated) │ 59 FXT science ObsIDs (PROVISIONAL) │ 215 Chandra catalog sources (179 within 8') │ P1 Algorithm tournament (FROZEN) │ 5 candidates: ILC · Poisson · Dictionary │ · Semi-blind · Sparse+Smooth │ 3 baselines: Broad-band · PS-mask · Hardness │ 72 scenarios × 5 metrics × pre-declared kill criteria │ → 8/8 INVALIDATED (expected at pilot stage) │ P2 Production XMM component cubes (FROZEN) │ exact-band per-ObsID/camera counts │ matched exposure / QPB / SP / spectral-line products │ counts closure + provenance closure validation │ P3 Production decomposition (FROZEN) │ real_data_ingestion → response_folding │ → production_ilc/poisson/dictionary/ │ semiblind/sparse → uncertainty_maps │ → production_compare │ wind/PS/sky/SP/QPB maps + uncertainty │ P4 XMM science measurements (4/10 in progress) │ discovery_measurement · null_test │ split_validation · real_epic_run │ + injection-recovery · Chandra cross-validation │ + closure · characterization · science report │ + freeze & audit │ P5 EP/FXT transfer (pending) │ FXT identifiability simulation + real-data results/limits │ P6 Publication/release (pending) code/tests/manifests/FITS/CSV/JSON/figures/Quarto

Acceptance criteria

The following thresholds were pre-declared before the tournament began and not modified after the verdicts:

Note: the kill thresholds used at the tournament stage were stricter (leakage ≤ 50%, flux_bias ≤ 50%, PA_bias ≤ 30°), and still invalidated all methods. The P4 stage will use the stricter scientific acceptance criteria above.

Current phase status

FROZEN P0 project/provenance reset V5 immutable fixed point · SHA-256 all-module pinning
FROZEN P1 algorithm tournament 8 methods × 72 scenarios · all INVALIDATED
FROZEN P2 production XMM component cubes exact-band counts + exposure/QPB/SP products
FROZEN P3 production decomposition real EPIC data + response folding + exposure correction
4/10 P4 XMM science measurements in progress · discovery/null/split/real_epic complete
PENDING P5 EP/FXT transfer FXT identifiability simulation + real-data application
PENDING P6 publication/release code/tests/manifests/FITS/Quarto release package

Test coverage: 90 test files, 2954 tests all passing.

XMM-Newton EPIC observation inventory

12 independent ObsIDs covering the M1 (14 files), M2 (12 files), PN (11 files) cameras:

0110930201 0147760101 0802710101 0802710201 0802710301 0802710401 0802710501 0802710601 0802710701 0802710801 0802710901 0802711001

Filter distribution: Medium (33 files) · Thin1 (2 files) · CalClosed (2 files, excluded). 23 complete ARF+RMF+PI response triplets passed OGIP-standard validation.

Next: P4 science measurements (Tasks 5–10)

  1. Injection-recovery: inject known wind signals at real-data noise levels and verify recovery accuracy and systematic bias. This is the gold standard for validating algorithm reliability.
  2. Chandra cross-validation: use Chandra's high-resolution point-source positions and inner-wind morphology as an independent cross-check of the XMM recovery. Chandra historical fluxes are not used as fixed constraints on XMM amplitudes.
  3. Closure tests: verify the complete loop raw counts → component decomposition → reprojection → comparison against the original data, ensuring no information is lost or introduced.
  4. Characterization: after Discovery mode is validated, use the recovered axis/cone families to sharpen PA, extent, asymmetry, and flux constraints.
  5. Science report: synthesize all measurements into the physical parameters of the NGC 3079 thermal wind: PA, opening angle, radial extent, N/S asymmetry, surface brightness profile, and component flux.
  6. Freeze & audit: freeze the P4 snapshot and launch an independent independent-agent-auditing review, ensuring all results are traceable and reproducible.

Provenance and immutability

P0 fixed point: P0_FIXED_POINT_V5_20260715T215301Z Manifest SHA-256: 3904ce… Metadata SHA-256: 2753f7… Ledger SHA-256: 8076c2… All input files: SHA-256 pinned All modules: hash verified V3/V4 rejected (precursor snapshots) Benchmark harness: benchmark-harness-v1 Compare engine: compare-v1 0 Critical/Important findings (V5 three-lane) 215 Chandra catalog sources (179 within 8') 59 FXT science ObsIDs (PROVISIONAL)

All results trace back to the immutable V5 fixed point. The provenance chain is complete: input inventory → SHA-256 pinning → module hash verification → audit ledger. Any result change must trace back to the final science goal, with the dated decision trail recorded in the cognition OS.

中文