diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index e073663..aaf74bf 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,12 +1,18 @@ # Contributing -Install `.[dev,viz]`, then run `pytest`, `ruff check .`, and `ruff format --check .`. -Keep simulator tests in a separate process from real-GPU checks. For numerical -changes, include an explicit expected result or independent reference. +Install `.[dev,viz]`; run `pytest`, `ruff check .`, and `ruff format --check .`. +See [tests/README.md](tests/README.md) for CPU, simulator, and hardware commands. -New example strategies need a documented device-function spec and hand-calculated -trade case. Keep demonstrations small and use synthetic data. Do not include -credentials, personal market datasets, private strategy plugins, or research output. +Keep boundaries clear: -Document numerical and execution-rule changes in the same pull request. New tracked -files must be listed in `scripts/check_release.py` after reviewing their contents. +- `src/gpu_backtest/core/` contains generic computation only. It must not import + examples, test references, benchmark code, CLI, workflows, or deployment tools. +- `workflows/` uses core primitives for splitting/analysis/output charts. +- Example trading rules/data stay under `examples/` and helper code under `tools/`. +- CLI modules are thin adapters; they do not own numerical or cloud logic. + +New example strategies need a device-function spec and hand-calculated case. +Use synthetic data. Keep credentials, private plugins/datasets, and research output +out of Git. New tracked files require a reviewed update to the explicit allowlist +in `scripts/check_release.py`. Preserve historical benchmark JSONs; new claims need +new measured evidence. Document intentional numerical/contract changes in the PR. diff --git a/MANIFEST.in b/MANIFEST.in index 7a9f8d3..5cf932e 100644 --- a/MANIFEST.in +++ b/MANIFEST.in @@ -1,9 +1,9 @@ include LICENSE README.md CONTRIBUTING.md recursive-include src/gpu_backtest *.py -recursive-include src/gpu_backtest *.txt -recursive-include tests *.py +recursive-include tools/gpu_backtest_tools *.py *.txt +recursive-include examples/gpu_backtest_examples *.py *.json *.csv *.md +recursive-include tests *.py *.md recursive-include docs *.md -recursive-include benchmarks *.json -recursive-include examples *.py *.json *.csv +recursive-include benchmarks *.json *.md recursive-include scripts *.py global-exclude __pycache__ *.py[cod] .env *.pem *.key diff --git a/README.md b/README.md index f5bcb75..b1cfd1d 100644 --- a/README.md +++ b/README.md @@ -2,131 +2,64 @@ ## 10× faster on our public billion-pair benchmark -**CPU: 7 min 44.60 s → RTX 4090: 44.95 s.** Measured on the same -**1,000,000,000-pair RSI grid × 1,024 bars**, against an **eight-thread compiled -Numba CPU baseline**. Exact measured speedup: **10.34×**; cloud setup is additional. +**CPU: 7 min 44.60 s → RTX 4090: 44.95 s.** Same RSI grid: +**1,000,000,000 pairs × 1,024 bars**, compared with an **eight-thread compiled +Numba CPU baseline**. Measured speedup **10.34×**, saving about seven minutes per +sweep. Cloud setup is additional; [method and raw evidence](docs/benchmarks.md). -**Backtest your own trading algorithms. Sweep a billion parameter combinations on a GPU.** +- **Your algorithm:** load a strategy module/object; private rules can remain private. +- **Large grids:** deterministic GPU reductions without materializing the full return matrix. +- **Your machine or RunPod:** use a local NVIDIA GPU or the separate cloud helper. -- **Swap algorithms:** load a separately installed strategy module or object. - Your strategy can stay in a private repo; the engine does not need to be edited. -- **10× faster — CPU minutes → GPU seconds:** the same **billion-pair RSI sweep** took - **7 min 44.60 s on an eight-thread Numba CPU baseline → 44.95 s on RTX 4090**. - That's **10.34× faster**, saving approximately **seven minutes per sweep**. -- **Billion-scale sweep:** **1,000,000,000 unique pairs × 1,024 bars**, with - both sides measured through statistics and CSV/manifest output. -- **No GPU in your laptop:** use the RunPod launcher to rent, run, download, and clean up. +## Where everything lives -This is a measured full-grid comparison, not an extrapolation from a tiny case. -See [the benchmark and raw data](docs/benchmarks.md) for hardware, timing scope, -and reproduction. Results depend on workload; cloud setup time is additional. - -The engine provides deterministic reductions and separate entry/exit effect-size -rankings without storing a full pairwise return matrix. This distribution includes -one educational RSI strategy and generated synthetic OHLCV data. Custom plugins -must follow [the Numba device-function contract](docs/strategy-contract.md). - -## Published performance - -| Public RSI workload | Compiled CPU, eight threads | RTX 4090 | Time saved | -|---|---|---|---| -| **1,000,000,000 pairs × 1,024 bars** | **7 min 44.60 s** | **44.95 s** | **6 min 59.65 s per sweep; 10.34× faster** | - -The CPU is an AMD EPYC 7K62 host running a parallel, compiled Numba baseline. -Both measurements use the same data, grid, fees, and two-pass algorithm, through -ranked CSV/manifest output. The CPU reducer and CUDA context were already warmed; -GPU kernel construction/JIT is included. Each full-grid timing is one measured run. -These are engine-job times, excluding pod provisioning and installation. - -**When GPU is useful:** repeated large parameter searches, where saving minutes -on every sweep adds up. If your job already finishes in a few CPU seconds, -renting/setup overhead may outweigh GPU savings. The smaller warmed-kernel timing -tests remain in the detailed benchmark, rather than being the main use-case claim. +```text +src/gpu_backtest/ + core/ GPU engine, kernels, grids, indicators, statistics, output + workflows/ Generic split/common analysis and charts + cli/ Thin command-line adapters +examples/gpu_backtest_examples/ + rsi/ Example strategy + config + generated CSV + data.py Example/benchmark synthetic data generator +tools/gpu_backtest_tools/ + runpod/ Optional API / SSH / bundle / lifecycle helper + benchmarks/ Performance runner and compiled CPU baseline + checks/ Hardware smoke checks and CPU test reference +tests/ + cpu/ CPU contracts, examples, workflows, helper tests + gpu/ Isolated CUDA simulation and real-GPU tests +docs/ Strategy contract, RunPod usage, benchmark method +benchmarks/results/ Historical measurements and validation evidence +``` -The billion grid uses 20,000 entry sets × 50,000 exit sets. Its grouped arrays -occupy 1.12 MB, compared with 4 GB for a full float32 return matrix; this excludes -input/indicator tables and runtime overhead. Pair counts are parameter combinations, -not trades. These measurements use public code and synthetic data, with no private -algorithm or market dataset. +**Start with `core/engine.py`** for the GPU run. Numerical kernels are in +`core/kernels.py`; trading rules are supplied by a plugin. Core imports no example, +CPU comparison engine, benchmark, or RunPod code. CPU preprocessing of market data +and indicator tables is part of the GPU pipeline; the separate CPU backtest baseline +is only a benchmark/reference tool. -## Quickstart without a GPU +## Try the RSI example without a GPU -Python 3.11–3.13 is supported. From a checkout: +Python 3.11–3.13, from a checkout: ```bash python3 -m venv .venv source .venv/bin/activate python -m pip install -e '.[dev,viz]' NUMBA_ENABLE_CUDASIM=1 gpu-backtest run \ - --config examples/rsi.json --out-prefix runs/rsi + --config examples/gpu_backtest_examples/rsi/config.json --out-prefix runs/rsi gpu-backtest charts --input runs/rsi_top_entry.csv runs/rsi_top_exit.csv \ --output runs/rsi.html ``` -Open `runs/rsi.html` in a browser. Its Vega libraries load from a public CDN. -The CUDA simulator is for tiny demonstrations and correctness tests. Large grids -need a real NVIDIA GPU. - -The example sweeps four entry combinations against four exit combinations. -It writes two ranked CSVs and an output manifest: - -```text -runs/rsi_top_entry.csv -runs/rsi_top_exit.csv -runs/rsi_top_manifest.json -``` - -`examples/synthetic.csv` contains generated bars, not historical market data. -Regenerate it with `python examples/generate_data.py`. - -## RunPod: use a GPU from your laptop - -The same engine runs on a rented RunPod GPU. With a funded account, Pod API key, -and registered SSH key: - -```bash -gpu-backtest runpod --config examples/rsi.json \ - --ssh-key ~/.ssh/runpod_ed25519 --output-dir runs/runpod-rsi --charts -``` - -This leases one RTX 4090, uploads selected engine/data/plugin files, installs the -environment, runs hardware numeric checks and the pipeline, downloads output, -and deletes the pod. Add `--dry-run` to inspect the bundle without renting anything; -`--mode run` selects one sweep and `--mode check` runs only hardware checks. -Private modules can be supplied with `--plugin-dir` without adding them to this repo. -See [the RunPod guide](docs/runpod.md) for setup, the manual route, environment -requirements, time/rate limits, custom plugins, and cleanup behavior. - -## NVIDIA GPU setup (inside the GPU machine) - -Run these commands in the GPU computer's terminal. For RunPod, use the pinned -setup in [the RunPod guide](docs/runpod.md); your laptop only controls the launcher. -For another GPU workstation/server, install a compatible driver and CUDA runtime/toolkit. The separate NVIDIA -`numba-cuda` backend uses the same `from numba import cuda` interface: - -```bash -python -m pip install -e '.[cuda]' -# If you also need CUDA 12 Python runtime/toolkit dependencies: -python -m pip install 'numba-cuda[cu12]>=0.30,<0.31' -gpu-backtest run --config examples/rsi.json --out-prefix runs/rsi_gpu -``` - -See the [NVIDIA installation guide](https://nvidia.github.io/numba-cuda/user/installation.html) -for CUDA 12/13 and driver requirements. Run GPU commands without -`NUMBA_ENABLE_CUDASIM=1`. The project's tested Numba series is 0.65; upgrading the -backend requires rerunning the simulator and real-GPU checks. +The simulator is for tiny examples/tests. Open `runs/rsi.html` in a browser; +its Vega libraries load from a public CDN. The [RSI example](examples/gpu_backtest_examples/rsi/README.md) +is educational and uses generated data. It is packaged separately as +`gpu_backtest_examples.rsi.strategy`, not as an engine builtin. ## Use your own strategy -Install the package containing your strategy in the same Python environment: - -```bash -python -m pip install -e /path/to/your-strategy-package -gpu-backtest run --strategy my_strategies.example \ - --input /path/to/market.csv --out-prefix runs/custom -``` - -Or pass a module/object through the Python API: +Install your strategy package into the same environment: ```python from gpu_backtest import run @@ -135,55 +68,44 @@ from my_strategies import example results = run(example, "market.csv", "runs/custom", buy=0.0015, sell=0.0015) ``` -The engine loads an installed Python module. Plugins are trusted code, and run -locally; they must implement the documented Numba device-function contract. -See [the strategy contract](docs/strategy-contract.md) for parameter specifications, -indicator tables, CPU references, and the RSI execution rules. +Or use `gpu-backtest run --strategy my_strategies.example --input market.csv +--out-prefix runs/custom`. Plugins implement a complete Numba CUDA trading loop; +see [the contract](docs/strategy.md). Entry/exit parameters are separate Cartesian +axes (up to four dimensions each), with up to four precomputed tables. Shared-parameter +diagonal-only sweeps are not supported. The API returns grouped sums/squares and +optional ranked artifacts, not a trade ledger, equity curve, Sharpe, or drawdown series. + +## Optional workflows and helpers -## Date splits and common parameters +| Command | Owner / purpose | +|---|---| +| `pipeline` | `workflows/`: split, run each segment, intersect ranked parameters | +| `split`, `common`, `charts` | `workflows/`: standalone data/result analysis | +| `runpod` | `tools/runpod/`: lease one GPU, upload selected files, download, delete | +| `benchmark` | `tools/benchmarks/`: reproduce the public CPU/GPU measurements | +| `gpu-check` | `tools/checks/`: small numeric checks on actual hardware | ```bash -NUMBA_ENABLE_CUDASIM=1 gpu-backtest pipeline \ - --config examples/rsi.json --output-dir runs/pipeline +gpu-backtest runpod --config examples/gpu_backtest_examples/rsi/config.json \ + --ssh-key ~/.ssh/runpod_ed25519 --output-dir runs/runpod-rsi --charts ``` -With three splits, the pipeline evaluates the full period and its two disjoint -halves. It records split hashes, writes ranked artifacts for every split, and -finds parameter combinations appearing in every top artifact. The output includes -`common_entry.csv` and `common_exit.csv` with per-period, average, and minimum -effect sizes. No common combinations is reported as a failed analysis, not an -empty successful result. - -Config `input` paths resolve relative to the JSON file. CLI `--input` and output -paths resolve relative to the current directory. A strategy must always be -specified explicitly. Config keys are `strategy`, `input`, `buy`, `sell`, -`entry_dims`, `exit_dims`, `top_n`, `threads_per_block`, `expected_interval`, -`num_splits`, `start_date`, `end_date`, and `ranges`. Dates use `DD-MM-YYYY`; -custom ranges are arrays such as `[["01-01-2024", "31-01-2024"], ...]`. - -`gpu-backtest split`, `common`, and `charts` also work independently; run each -subcommand with `--help` for its arguments. - -## Scope and interpretation - -- Entry and exit parameters form a Cartesian product, with at most four - dimensions on each side and four precomputed indicator tables. Shared-parameter - diagonal-only sweeps are not supported. -- The engine runs two passes over the same return matrix. Returns and squared - returns are rounded to float32, then accumulated in float64 in a fixed order. -- Ranked rows represent entry/exit parameter groups, not individually selected - complete strategies. Their observations are parameter combinations, not - independent market samples. Effect sizes and nominal normal intervals are - descriptive grid comparisons. `posterior_prob_superior` is a normal-CDF proxy; - it is not a Bayesian probability or a forecast of profitable trading. -- Strategy plugins own their trading loop, position sizing, and execution rules. - The included example uses long-only next-open execution, with its boundary and - fee conventions documented in the strategy contract. -- The current API returns grouped sums and sums of squares, plus optional top - artifacts. It does not return a trade ledger, equity curve, Sharpe ratio, or - drawdown series. Example performance is not a trading recommendation. - -## Development +RunPod requires your account/API key and registered SSH key. It is an optional +execution helper; [setup, manual GPU route, and cleanup](docs/runpod.md). +Commands keep their existing names. The public `from gpu_backtest import run` API +and old `rsi_meanrev` shorthand remain usable; direct internal imports moved under +`core/`, `workflows/`, or `gpu_backtest_tools` in v0.5. + +## Performance and development + +| Same billion-pair RSI job | CPU, eight threads | RTX 4090 | Saved per sweep | +|---|---|---|---| +| 1,000,000,000 pairs × 1,024 bars, through output | 7 min 44.60 s | 44.95 s | 6 min 59.65 s; 10.34× faster | + +This is one measured run per side on the published setup, not a universal speed +claim. Small CPU jobs may not justify cloud startup. Statistics describe parameter +combination groups, not independent market samples or a forecast of profits. +See [benchmark details](docs/benchmarks.md) for scope and raw data. ```bash python -m pytest @@ -193,21 +115,7 @@ python -m build python scripts/check_release.py ``` -Default tests need no GPU. They run simulation in a separate process and verify -hand-calculated results, known reduction matrices, independent CPU references, -output determinism, statistics, indicator values, split coverage, and external -plugin loading. For a real NVIDIA GPU: - -```bash -NUMBA_ENABLE_CUDASIM=0 python -m pytest -m gpu -v -``` - -See [testing details](docs/testing.md) and the [RunPod validation record](docs/runpod-validation.md). -Small numeric checks and the public RSI pipeline have passed on a real RTX 4090. -The [public benchmark](docs/benchmarks.md) also completed a billion-pair RSI sweep. -Other hardware/backend versions and custom strategies require their own validation. - -## License +[CPU/GPU testing](tests/README.md) · [Plugin contract](docs/strategy.md) · +[RunPod helper](docs/runpod.md) · [Benchmark method](docs/benchmarks.md) -[MIT](LICENSE). You may use the engine with separately maintained private -strategy plugins, subject to their own licenses. +[MIT](LICENSE). Private strategy plugins retain their own licenses. diff --git a/docs/runpod-validation.md b/benchmarks/results/runpod_validation_20261003.md similarity index 96% rename from docs/runpod-validation.md rename to benchmarks/results/runpod_validation_20261003.md index c912c32..44201ae 100644 --- a/docs/runpod-validation.md +++ b/benchmarks/results/runpod_validation_20261003.md @@ -20,7 +20,7 @@ confirmed it was absent from the account's pod list. | Optional HTML renderer | Altair 6.3.0 | The launcher now pins the resolved core packages in -[`runpod_requirements.txt`](../src/gpu_backtest/runpod_requirements.txt). +[`runpod_requirements.txt`](../../tools/gpu_backtest_tools/runpod/requirements.txt). ## Results diff --git a/docs/runpod.md b/docs/runpod.md index 26cf525..834240d 100644 --- a/docs/runpod.md +++ b/docs/runpod.md @@ -31,7 +31,7 @@ remote job, and state file. Server error bodies/auth headers are not printed. From an installed checkout: ```bash -gpu-backtest runpod --config examples/rsi.json \ +gpu-backtest runpod --config examples/gpu_backtest_examples/rsi/config.json \ --ssh-key ~/.ssh/runpod_ed25519 --output-dir runs/runpod-rsi --charts ``` @@ -48,7 +48,7 @@ The output directory must be new or empty. A dry run reads no API credentials, requires no SSH key, and rents no GPU: ```bash -gpu-backtest runpod --config examples/rsi.json \ +gpu-backtest runpod --config examples/gpu_backtest_examples/rsi/config.json \ --output-dir runs/runpod-preview --dry-run ``` @@ -97,7 +97,7 @@ runpod/pytorch:2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04 This existing official image supplies Python 3.11, SSH, and CUDA development libraries. The engine does not use its PyTorch installation. The launcher creates an isolated venv and installs the engine with the pinned -[validated environment](../src/gpu_backtest/runpod_requirements.txt), using the +[validated environment](../tools/gpu_backtest_tools/runpod/requirements.txt), using the image's CUDA toolkit. Optional charts use Altair 6.3.0. Overrides must supply Python 3.11–3.13, compatible CUDA development libraries/driver, SSH, tar, and `nvidia-smi`. @@ -151,7 +151,7 @@ source .venv/bin/activate python -m pip install -e '.[dev,viz,cuda]' 'cuda-bindings>=12.9.1,<13' NUMBA_ENABLE_CUDASIM=0 gpu-backtest gpu-check --output runs/gpu_check.json NUMBA_ENABLE_CUDASIM=0 gpu-backtest pipeline \ - --config examples/rsi.json --output-dir runs/example + --config examples/gpu_backtest_examples/rsi/config.json --output-dir runs/example gpu-backtest charts --input runs/example/*_top_entry.csv runs/example/*_top_exit.csv \ --output runs/example/charts.html ``` @@ -161,5 +161,5 @@ launcher deletes only pods it creates. Small hardware checks establish GPU compilation and explicit numeric cases, not large-grid speed or every strategy's correctness. -See [the hardware validation record](runpod-validation.md) for the tested image, +See [the hardware validation record](../benchmarks/results/runpod_validation_20261003.md) for the tested image, driver, package versions, numeric results, and successful cleanup. diff --git a/docs/strategy-contract.md b/docs/strategy.md similarity index 69% rename from docs/strategy-contract.md rename to docs/strategy.md index 164e050..8b4150c 100644 --- a/docs/strategy-contract.md +++ b/docs/strategy.md @@ -1,7 +1,9 @@ # Strategy contract -Pass a module/object to `gpu_backtest.run`, a builtin name such as `rsi_meanrev`, -or an installed dotted module name such as `my_strategies.example`. +Pass a module/object to `gpu_backtest.run`, or an installed dotted module name +such as `my_strategies.example`. The separate example module is +`gpu_backtest_examples.rsi.strategy`. The old `rsi_meanrev` shorthand is translated +only at the public API/CLI edge; the core imports no example strategy. The package neither copies nor uploads an external plugin. ## Required declarations @@ -89,27 +91,11 @@ def reference( ): ... # Independent CPU implementation returning percent return. ``` -`gpu_backtest.engine.reference_group_sums` uses this reference on small grids. +`gpu_backtest_tools.checks.reference.reference_group_sums` uses this reference on small grids. Supply a hand-calculated case as well: matching two implementations alone does not establish that the intended rules are correct. -## Included RSI example execution rules +## Example -Entry: RSI(`p_e`) below `buy_lvl` while flat. Exit: RSI(`p_x`) above `sell_lvl` -while holding a position. The initial capital is 10,000 units. - -Signals are evaluated after processing pending orders, starting at bar 1. -Orders fill at the following bar's open, with exits processed before entries. -The loop processes bars 0 through `n-2`. Pending orders are **not filled on the -final bar**; an existing position is valued at that bar's open, with no forced -sale or final sell fee. This boundary convention is covered by tests. - -Buying invests all current capital before charging the buy fee. Cash therefore -becomes negative by the fee amount; this example permits that financing rather -than reserving the fee within available cash. Selling credits proceeds minus -the sell fee. There is no slippage, funding charge, leverage model, or partial fill. - -For the seven-bar hand-calculated fixture, the zero-fee trade buys 125 units at -80 and sells them at 120: return 50%. At 0.15% per side, buy fee is 15 and sell -fee is 22.5; final equity is 14,962.5, giving 49.625%. Tests also cover a losing -round trip, no-entry conditions, and ignored final-bar pending orders. +See [the separate RSI example](../examples/gpu_backtest_examples/rsi/README.md) +for a complete plugin, runnable config/data, and its execution/fee conventions. diff --git a/docs/testing.md b/docs/testing.md deleted file mode 100644 index 8f2a667..0000000 --- a/docs/testing.md +++ /dev/null @@ -1,63 +0,0 @@ -# Testing and release checks - -The default pytest command runs CPU checks and a child process with -`NUMBA_ENABLE_CUDASIM=1`. Simulation is isolated from the parent so that it cannot -silently turn a later GPU run into CPU simulation. - -The tests use generated OHLCV data and explicit numeric fixtures. No downloaded -market data, external credentials, or cloud resources are required. - -## Proofs used - -- RSI hand calculations anchor timing, fee accounting, a losing trade, and - last-bar mark-to-market behavior independently of implementation agreement. -- A tiny algebraic plugin has a known return matrix. Both GPU reduction passes - must match exact row/column sums and float32 sums of squares. The plugin is - created outside the installed engine, exercising the external module path. -- RSI GPU-device results on a synthetic grid are compared with the CPU reference. -- Repeating the same simulated sweep must produce identical arrays and byte-identical - top CSVs/manifests. This is a regression/determinism check, not an independent - proof of strategy correctness. -- Group statistics are compared with explicit group-versus-rest calculations. -- Indicators have hand-calculated spot values. Splits must be disjoint and cover - the requested range; common artifacts preserve decimal precision and deterministic - tie ordering. -- CLI run/pipeline and optional HTML charts are tested with public sample inputs. -- Plugin validation rejects malformed dimensions, reordered overrides, invalid - table references, oversized specs, and reserved output names before device work. - -## Real GPU - -```bash -NUMBA_ENABLE_CUDASIM=0 python -m pytest -m gpu -v -``` - -GPU tests fail if simulation is enabled, and skip if no NVIDIA GPU is available. -Require actual passes, rather than skips, before claiming hardware validation. -The packaged `gpu-backtest gpu-check` command performs numeric hardware checks -without pytest. It and the public RSI pipeline passed on an RTX 4090; see the -[validation record](runpod-validation.md). This establishes small-grid GPU -compilation/results, not large-grid performance or arbitrary plugin correctness. - -RunPod lifecycle tests simulate create/SSH/job/download/cleanup failures, timeouts, -normal cancellation, cancellation during creation, price rejection, ambiguous -allocation responses, metadata-write failure, upload selection, and unsafe tar -entries. They do not rent GPUs during default tests. - -The [public performance benchmark](benchmarks.md) separately measures warmed -compiled CPU/GPU reductions and a normal billion-pair engine run. Its CPU baseline -is checked against the Python reference, including float32 sums of squares; -default tests verify the billion grid size without executing that grid. -`--billion-cpu` additionally runs the full billion-pair CPU job through output and -checks all four aggregate arrays against the GPU run; it never substitutes an estimate. - -## Distribution contents - -`scripts/check_release.py` checks an explicit tracked-file allowlist and common -credential patterns. CI runs it and builds a source distribution and wheel. -The public tree contains the engine, one educational RSI plugin, generated data, -documentation, and generic tests. New files require an explicit allowlist update. - -Before publishing a release, inspect both distribution archives, execute the -documented example from a clean installation, and review the tracked diff. The -allowlist complements manual content review; it is not a proof against all secrets. diff --git a/examples/generate_data.py b/examples/generate_data.py deleted file mode 100644 index 8cfdddc..0000000 --- a/examples/generate_data.py +++ /dev/null @@ -1,34 +0,0 @@ -"""Generate deterministic synthetic daily OHLCV bars; no market data is used.""" - -import csv -import math -from datetime import datetime, timedelta, timezone -from pathlib import Path - - -def generate(output, bars=128): - output = Path(output) - output.parent.mkdir(parents=True, exist_ok=True) - previous = 100.0 - with output.open("w", newline="") as handle: - writer = csv.writer(handle, lineterminator="\n") - writer.writerow(["Time", "Open", "High", "Low", "Close", "Volume"]) - for i in range(bars): - close = 100 + 16 * math.sin(i * 0.45) + 4 * math.sin(i * 1.1) + i * 0.03 - open_ = previous + 0.4 * math.cos(i) - high = max(open_, close) + 1.0 - low = min(open_, close) - 1.0 - time = datetime(2024, 1, 1, tzinfo=timezone.utc) + timedelta(days=i) - writer.writerow( - [ - time.isoformat(), - *[f"{v:.6f}" for v in (open_, high, low, close)], - 1000 + 50 * (i % 7), - ] - ) - previous = close - return output - - -if __name__ == "__main__": - print(generate(Path(__file__).with_name("synthetic.csv"))) diff --git a/examples/gpu_backtest_examples/__init__.py b/examples/gpu_backtest_examples/__init__.py new file mode 100644 index 0000000..ded36c4 --- /dev/null +++ b/examples/gpu_backtest_examples/__init__.py @@ -0,0 +1 @@ +"""Educational examples; the engine core never imports this package.""" diff --git a/src/gpu_backtest/synthetic.py b/examples/gpu_backtest_examples/data.py similarity index 100% rename from src/gpu_backtest/synthetic.py rename to examples/gpu_backtest_examples/data.py diff --git a/examples/gpu_backtest_examples/rsi/README.md b/examples/gpu_backtest_examples/rsi/README.md new file mode 100644 index 0000000..1ca0d1d --- /dev/null +++ b/examples/gpu_backtest_examples/rsi/README.md @@ -0,0 +1,41 @@ +# RSI example plugin + +This example lives outside `gpu_backtest.core`. Entry is RSI below `buy_lvl`; +exit is RSI above `sell_lvl`. Entry/exit periods are independently swept. + +| File | Purpose | +|---|---| +| `strategy.py` | Parameter specs, GPU device function, and CPU reference | +| `config.json` | A tiny runnable grid and fees | +| `synthetic.csv` | Generated data, not historical market data | +| `generate_data.py` | Regenerate the fixture via the shared example data generator | + +```bash +NUMBA_ENABLE_CUDASIM=1 gpu-backtest run \ + --config examples/gpu_backtest_examples/rsi/config.json --out-prefix runs/rsi +python -m gpu_backtest_examples.rsi.generate_data +``` + +The installed module is `gpu_backtest_examples.rsi.strategy`. The engine loads +it through the same contract used by an external private strategy. + +## Execution rules + +Entry: RSI(`p_e`) below `buy_lvl` while flat. Exit: RSI(`p_x`) above `sell_lvl` +while holding a position. The initial capital is 10,000 units. + +Signals are evaluated after processing pending orders, starting at bar 1. +Orders fill at the following bar's open, with exits processed before entries. +The loop processes bars 0 through `n-2`. Pending orders are **not filled on the +final bar**; an existing position is valued at that bar's open, with no forced +sale or final sell fee. This boundary convention is covered by tests. + +Buying invests all current capital before charging the buy fee. Cash therefore +becomes negative by the fee amount; this example permits that financing rather +than reserving the fee within available cash. Selling credits proceeds minus +the sell fee. There is no slippage, funding charge, leverage model, or partial fill. + +For the seven-bar hand-calculated fixture, the zero-fee trade buys 125 units at +80 and sells them at 120: return 50%. At 0.15% per side, buy fee is 15 and sell +fee is 22.5; final equity is 14,962.5, giving 49.625%. Tests also cover a losing +round trip, no-entry conditions, and ignored final-bar pending orders. diff --git a/examples/gpu_backtest_examples/rsi/__init__.py b/examples/gpu_backtest_examples/rsi/__init__.py new file mode 100644 index 0000000..e7d5677 --- /dev/null +++ b/examples/gpu_backtest_examples/rsi/__init__.py @@ -0,0 +1 @@ +"""Example RSI strategy and its runnable config/data.""" diff --git a/examples/gpu_backtest_examples/rsi/config.json b/examples/gpu_backtest_examples/rsi/config.json new file mode 100644 index 0000000..44b12c4 --- /dev/null +++ b/examples/gpu_backtest_examples/rsi/config.json @@ -0,0 +1,41 @@ +{ + "strategy": "gpu_backtest_examples.rsi.strategy", + "input": "synthetic.csv", + "buy": 0.0015, + "sell": 0.0015, + "entry_dims": [ + [ + "p_e", + 2, + 4, + 2, + false + ], + [ + "buy_lvl", + 20, + 30, + 10, + false + ] + ], + "exit_dims": [ + [ + "p_x", + 2, + 4, + 2, + false + ], + [ + "sell_lvl", + 70, + 80, + 10, + false + ] + ], + "top_n": 4, + "expected_interval": "1D", + "num_splits": 3 +} diff --git a/examples/gpu_backtest_examples/rsi/generate_data.py b/examples/gpu_backtest_examples/rsi/generate_data.py new file mode 100644 index 0000000..f37c85b --- /dev/null +++ b/examples/gpu_backtest_examples/rsi/generate_data.py @@ -0,0 +1,8 @@ +"""Regenerate only the public synthetic RSI example.""" + +from pathlib import Path + +from gpu_backtest_examples.data import generate_csv + +if __name__ == "__main__": + print(generate_csv(Path(__file__).with_name("synthetic.csv"))) diff --git a/src/gpu_backtest/strategies/rsi_meanrev.py b/examples/gpu_backtest_examples/rsi/strategy.py similarity index 100% rename from src/gpu_backtest/strategies/rsi_meanrev.py rename to examples/gpu_backtest_examples/rsi/strategy.py diff --git a/examples/synthetic.csv b/examples/gpu_backtest_examples/rsi/synthetic.csv similarity index 100% rename from examples/synthetic.csv rename to examples/gpu_backtest_examples/rsi/synthetic.csv diff --git a/examples/rsi.json b/examples/rsi.json deleted file mode 100644 index 309a2f3..0000000 --- a/examples/rsi.json +++ /dev/null @@ -1,11 +0,0 @@ -{ - "strategy": "rsi_meanrev", - "input": "synthetic.csv", - "buy": 0.0015, - "sell": 0.0015, - "entry_dims": [["p_e", 2, 4, 2, false], ["buy_lvl", 20, 30, 10, false]], - "exit_dims": [["p_x", 2, 4, 2, false], ["sell_lvl", 70, 80, 10, false]], - "top_n": 4, - "expected_interval": "1D", - "num_splits": 3 -} diff --git a/pyproject.toml b/pyproject.toml index 5b4f4a9..ac5c488 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "gpu-backtest-engine" -version = "0.4.0" +version = "0.5.0" description = "Pluggable GPU parameter sweeps with deterministic entry/exit effect-size analysis" readme = "README.md" requires-python = ">=3.11,<3.14" @@ -25,11 +25,19 @@ gpu-backtest = "gpu_backtest.cli:main" Repository = "https://github.com/howardc38/gpu-backtest-engine" Issues = "https://github.com/howardc38/gpu-backtest-engine/issues" +[tool.setuptools.package-dir] +gpu_backtest = "src/gpu_backtest" +gpu_backtest_examples = "examples/gpu_backtest_examples" +gpu_backtest_tools = "tools/gpu_backtest_tools" + [tool.setuptools.packages.find] -where = ["src"] +where = ["src", "examples", "tools"] +include = ["gpu_backtest", "gpu_backtest.*", "gpu_backtest_examples", "gpu_backtest_examples.*", "gpu_backtest_tools", "gpu_backtest_tools.*"] +namespaces = false [tool.setuptools.package-data] -gpu_backtest = ["runpod_requirements.txt"] +"gpu_backtest_tools.runpod" = ["requirements.txt"] +"gpu_backtest_examples.rsi" = ["config.json", "synthetic.csv"] [tool.pytest.ini_options] testpaths = ["tests"] diff --git a/scripts/check_release.py b/scripts/check_release.py index 513f28f..eac3dbe 100644 --- a/scripts/check_release.py +++ b/scripts/check_release.py @@ -13,58 +13,80 @@ "LICENSE", "MANIFEST.in", "README.md", - "pyproject.toml", - "docs/strategy-contract.md", - "docs/testing.md", - "docs/runpod.md", - "docs/runpod-validation.md", - "docs/benchmarks.md", "benchmarks/results/rtx4090_rsi_20261003.json", "benchmarks/results/rtx4090_rsi_matched_billion_20261003.json", - "examples/generate_data.py", - "examples/rsi.json", - "examples/synthetic.csv", + "benchmarks/results/runpod_validation_20261003.md", + "docs/benchmarks.md", + "docs/runpod.md", + "docs/strategy.md", + "examples/gpu_backtest_examples/__init__.py", + "examples/gpu_backtest_examples/data.py", + "examples/gpu_backtest_examples/rsi/README.md", + "examples/gpu_backtest_examples/rsi/__init__.py", + "examples/gpu_backtest_examples/rsi/config.json", + "examples/gpu_backtest_examples/rsi/generate_data.py", + "examples/gpu_backtest_examples/rsi/strategy.py", + "examples/gpu_backtest_examples/rsi/synthetic.csv", + "pyproject.toml", "scripts/check_release.py", "src/gpu_backtest/__init__.py", "src/gpu_backtest/__main__.py", - "src/gpu_backtest/benchmark.py", - "src/gpu_backtest/benchmark_cpu.py", - "src/gpu_backtest/charts.py", - "src/gpu_backtest/cli.py", - "src/gpu_backtest/common.py", - "src/gpu_backtest/engine.py", - "src/gpu_backtest/gpu_check.py", - "src/gpu_backtest/runpod.py", - "src/gpu_backtest/runpod_requirements.txt", - "src/gpu_backtest/pipeline.py", - "src/gpu_backtest/splits.py", - "src/gpu_backtest/statistics.py", - "src/gpu_backtest/strategy.py", - "src/gpu_backtest/synthetic.py", - "src/gpu_backtest/tables.py", - "src/gpu_backtest/strategies/__init__.py", - "src/gpu_backtest/strategies/rsi_meanrev.py", - "tests/helpers.py", - "tests/simulator_cases.py", - "tests/test_cli.py", - "tests/test_benchmark.py", - "tests/test_common_contract.py", - "tests/test_engine_contract.py", - "tests/test_gpu.py", - "tests/test_gpu_check.py", - "tests/test_runpod.py", - "tests/test_simulator.py", - "tests/test_split_contract.py", - "tests/test_statistics.py", - "tests/test_strategy.py", - "tests/test_tables.py", + "src/gpu_backtest/cli/__init__.py", + "src/gpu_backtest/cli/analysis.py", + "src/gpu_backtest/cli/backtest.py", + "src/gpu_backtest/cli/options.py", + "src/gpu_backtest/cli/tools.py", + "src/gpu_backtest/core/__init__.py", + "src/gpu_backtest/core/data.py", + "src/gpu_backtest/core/engine.py", + "src/gpu_backtest/core/grid.py", + "src/gpu_backtest/core/indicators.py", + "src/gpu_backtest/core/kernels.py", + "src/gpu_backtest/core/output.py", + "src/gpu_backtest/core/statistics.py", + "src/gpu_backtest/core/strategy.py", + "src/gpu_backtest/workflows/__init__.py", + "src/gpu_backtest/workflows/charts.py", + "src/gpu_backtest/workflows/common.py", + "src/gpu_backtest/workflows/pipeline.py", + "src/gpu_backtest/workflows/splits.py", + "tests/README.md", + "tests/cpu/test_architecture.py", + "tests/cpu/test_benchmark.py", + "tests/cpu/test_cli.py", + "tests/cpu/test_common_contract.py", + "tests/cpu/test_engine_contract.py", + "tests/cpu/test_examples.py", + "tests/cpu/test_gpu_check.py", + "tests/cpu/test_runpod.py", + "tests/cpu/test_split_contract.py", + "tests/cpu/test_statistics.py", + "tests/cpu/test_strategy.py", + "tests/cpu/test_tables.py", + "tests/gpu/fixtures.py", + "tests/gpu/simulator_cases.py", + "tests/gpu/test_hardware.py", + "tests/gpu/test_simulator.py", + "tools/gpu_backtest_tools/__init__.py", + "tools/gpu_backtest_tools/benchmarks/__init__.py", + "tools/gpu_backtest_tools/benchmarks/cpu.py", + "tools/gpu_backtest_tools/benchmarks/runner.py", + "tools/gpu_backtest_tools/checks/__init__.py", + "tools/gpu_backtest_tools/checks/gpu.py", + "tools/gpu_backtest_tools/checks/reference.py", + "tools/gpu_backtest_tools/runpod/__init__.py", + "tools/gpu_backtest_tools/runpod/bundle.py", + "tools/gpu_backtest_tools/runpod/client.py", + "tools/gpu_backtest_tools/runpod/launcher.py", + "tools/gpu_backtest_tools/runpod/requirements.txt", + "tools/gpu_backtest_tools/runpod/transport.py", } PATTERNS = [ - re.compile(r"AKIA[A-Z0-9]{16}"), - re.compile(r"gh[pousr]_[A-Za-z0-9]{30,}"), - re.compile(r"github_pat_[A-Za-z0-9_]{30,}"), - re.compile(r"-----BEGIN (?:RSA |EC |OPENSSH )?PRIVATE KEY-----"), - re.compile(r"(?:api_key|password|secret)\s*=\s*['\"][^'\"]{12,}['\"]", re.IGNORECASE), + re.compile("AKIA[A-Z0-9]{16}"), + re.compile("gh[pousr]_[A-Za-z0-9]{30,}"), + re.compile("github_pat_[A-Za-z0-9_]{30,}"), + re.compile("-----BEGIN (?:RSA |EC |OPENSSH )?PRIVATE KEY-----"), + re.compile("(?:api_key|password|secret)\\s*=\\s*['\\\"][^'\\\"]{12,}['\\\"]", re.IGNORECASE), ] @@ -72,7 +94,7 @@ def main(): result = subprocess.run(["git", "ls-files", "-z"], cwd=ROOT, capture_output=True) if result.returncode: raise SystemExit("Initialize and stage the new repository before checking release files") - paths = set(result.stdout.decode().split("\0")) - {""} + paths = set(result.stdout.decode().split("\x00")) - {""} extra = paths - ALLOWED missing = ALLOWED - paths if extra or missing: @@ -86,7 +108,7 @@ def main(): problems.append(f"{relative}: symlinks require explicit review") continue text = path.read_text(encoding="utf-8") - if any(pattern.search(text) for pattern in PATTERNS): + if any((pattern.search(text) for pattern in PATTERNS)): problems.append(f"{relative}: possible credential detected") if problems: raise SystemExit("\n".join(problems)) diff --git a/src/gpu_backtest/__init__.py b/src/gpu_backtest/__init__.py index a136da0..026173e 100644 --- a/src/gpu_backtest/__init__.py +++ b/src/gpu_backtest/__init__.py @@ -1,11 +1,26 @@ """Pluggable GPU parameter sweeps and grouped-return analysis.""" -__version__ = "0.4.0" +__version__ = "0.5.0" + + +def resolve_strategy_name(value): + """Keep the old demo shorthand at the API edge, outside the engine core.""" + if isinstance(value, str) and value in ( + "rsi_meanrev", + "gpu_backtest.strategies.rsi_meanrev", + ): + return "gpu_backtest_examples.rsi.strategy" + return value def run(*args, **kwargs): - """Run a strategy module or object; see gpu_backtest.engine.run.""" - from .engine import run as run_engine + """Run a strategy module or object; see gpu_backtest.core.engine.run.""" + from .core.engine import run as run_engine + + if args: + args = (resolve_strategy_name(args[0]), *args[1:]) + elif "strategy_name" in kwargs: + kwargs["strategy_name"] = resolve_strategy_name(kwargs["strategy_name"]) return run_engine(*args, **kwargs) diff --git a/src/gpu_backtest/cli.py b/src/gpu_backtest/cli.py deleted file mode 100644 index ff9f62e..0000000 --- a/src/gpu_backtest/cli.py +++ /dev/null @@ -1,179 +0,0 @@ -"""Command-line entry point with explicit strategy selection and JSON configuration.""" - -import argparse -import json -from pathlib import Path - - -def _options(args): - config = {} - if args.config: - config_path = Path(args.config).resolve() - config = json.loads(config_path.read_text()) - if not isinstance(config, dict): - raise ValueError("Config must be a JSON object") - if config.get("input"): - config["input"] = str(config_path.parent / config["input"]) - strategy = args.strategy if args.strategy is not None else config.get("strategy") - source = args.input if args.input is not None else config.get("input") - if not strategy or not source: - raise ValueError("Specify strategy and input explicitly, either in config or CLI flags") - options = {} - for name, default in ( - ("buy", 0.0015), - ("sell", 0.0015), - ("top_n", 100), - ("threads_per_block", 128), - ("expected_interval", None), - ): - value = getattr(args, name) - options[name] = config.get(name, default) if value is None else value - options.update(entry_dims=config.get("entry_dims"), exit_dims=config.get("exit_dims")) - return strategy, source, options, config - - -def main(argv=None): - parser = argparse.ArgumentParser(prog="gpu-backtest") - subparsers = parser.add_subparsers(dest="command", required=True) - for name in ("run", "pipeline"): - command = subparsers.add_parser(name) - command.add_argument("--config") - command.add_argument("--strategy", help="Builtin name or installed dotted module name") - command.add_argument("--input") - command.add_argument("--buy", type=float) - command.add_argument("--sell", type=float) - command.add_argument("--top-n", dest="top_n", type=int) - command.add_argument("--threads-per-block", dest="threads_per_block", type=int) - command.add_argument("--expected-interval", dest="expected_interval") - if name == "run": - command.add_argument("--out-prefix", required=True) - else: - command.add_argument("--output-dir", required=True) - split = subparsers.add_parser("split") - split.add_argument("--input", required=True) - split.add_argument("--start-date") - split.add_argument("--end-date") - split.add_argument("--num-splits", type=int, default=3) - common = subparsers.add_parser("common") - common.add_argument("--input", nargs="+", required=True) - common.add_argument("--keys", required=True) - common.add_argument("--labels") - common.add_argument("--output", required=True) - charts = subparsers.add_parser("charts") - charts.add_argument("--input", nargs="+", required=True) - charts.add_argument("--output", required=True) - gpu_check = subparsers.add_parser("gpu-check", help="Run small numeric checks on a real GPU") - gpu_check.add_argument("--output") - benchmark = subparsers.add_parser("benchmark", help="Measure compiled CPU/GPU RSI sweeps") - benchmark.add_argument("--output", required=True) - benchmark.add_argument("--bars", type=int, default=1024) - benchmark.add_argument("--cpu-threads", type=int, default=8) - benchmark.add_argument("--repeats", type=int, default=3) - benchmark.add_argument("--billion", action="store_true") - benchmark.add_argument( - "--billion-cpu", action="store_true", help="Measure the entire billion grid on CPU as well" - ) - runpod = subparsers.add_parser( - "runpod", help="Lease one GPU, run a selected job, download, clean up" - ) - runpod.add_argument("--config") - runpod.add_argument("--output-dir", required=True) - runpod.add_argument("--ssh-key") - runpod.add_argument("--plugin-dir") - runpod.add_argument( - "--mode", choices=("run", "pipeline", "check", "benchmark"), default="pipeline" - ) - from .runpod import DEFAULT_IMAGE - - runpod.add_argument("--image", default=DEFAULT_IMAGE) - runpod.add_argument("--gpu", default="NVIDIA GeForce RTX 4090") - runpod.add_argument("--cloud", choices=("SECURE", "COMMUNITY"), default="SECURE") - runpod.add_argument("--max-seconds", type=int, default=1800) - runpod.add_argument("--max-hourly-rate", type=float, default=1.0) - runpod.add_argument("--keep-pod", action="store_true") - runpod.add_argument("--charts", action="store_true") - runpod.add_argument("--dry-run", action="store_true") - args = parser.parse_args(argv) - try: - if args.command in ("run", "pipeline"): - strategy, source, options, config = _options(args) - if args.command == "run": - from .engine import run - - run(strategy, source, args.out_prefix, **options) - else: - from .pipeline import run_pipeline - - run_pipeline( - strategy, - source, - args.output_dir, - num_splits=config.get("num_splits", 3), - start_date=config.get("start_date"), - end_date=config.get("end_date"), - ranges=config.get("ranges"), - **options, - ) - elif args.command == "split": - from .splits import _read_input, split_csv_by_date_range - - data, column = _read_input(args.input, "Time") - split_csv_by_date_range( - args.input, - args.start_date or data[column].min().strftime("%d-%m-%Y"), - args.end_date or data[column].max().strftime("%d-%m-%Y"), - args.num_splits, - column, - ) - elif args.command == "common": - from .common import process_data - - process_data( - { - "input_files": args.input, - "key_columns": [p.strip() for p in args.keys.split(",") if p.strip()], - "file_labels": args.labels.split(",") if args.labels else None, - "output_file": args.output, - } - ) - elif args.command == "charts": - from .charts import write_charts - - print(write_charts(args.input, args.output)) - elif args.command == "gpu-check": - from .gpu_check import check_gpu - - check_gpu(args.output) - elif args.command == "benchmark": - from .benchmark import benchmark - - benchmark( - args.output, - bars=args.bars, - cpu_threads=args.cpu_threads, - repeats=args.repeats, - billion=args.billion, - billion_cpu=args.billion_cpu, - ) - else: - from .runpod import launch, terminate_on_signal - - with terminate_on_signal(): - launch( - config_path=args.config, - output_dir=args.output_dir, - ssh_key=args.ssh_key, - mode=args.mode, - plugin_dir=args.plugin_dir, - image=args.image, - gpu=args.gpu, - cloud=args.cloud, - max_seconds=args.max_seconds, - max_hourly_rate=args.max_hourly_rate, - keep_pod=args.keep_pod, - charts=args.charts, - dry_run=args.dry_run, - ) - except (ValueError, ImportError, OSError, KeyError, TypeError, RuntimeError) as exc: - parser.exit(2, f"gpu-backtest: {exc}\n") - return 0 diff --git a/src/gpu_backtest/cli/__init__.py b/src/gpu_backtest/cli/__init__.py new file mode 100644 index 0000000..ca8aa57 --- /dev/null +++ b/src/gpu_backtest/cli/__init__.py @@ -0,0 +1,18 @@ +"""Command-line adapters for the engine and optional tools.""" + +import argparse + +from . import analysis, backtest, tools + + +def main(argv=None): + parser = argparse.ArgumentParser(prog="gpu-backtest") + subparsers = parser.add_subparsers(dest="command", required=True) + for group in (backtest, analysis, tools): + group.register(subparsers) + args = parser.parse_args(argv) + try: + args.handler(args) + except (ValueError, ImportError, OSError, KeyError, TypeError, RuntimeError) as exc: + parser.exit(2, f"gpu-backtest: {exc}\n") + return 0 diff --git a/src/gpu_backtest/cli/analysis.py b/src/gpu_backtest/cli/analysis.py new file mode 100644 index 0000000..26e1cfc --- /dev/null +++ b/src/gpu_backtest/cli/analysis.py @@ -0,0 +1,49 @@ +"""Analysis commands consume generic input/results rather than strategy code.""" + + +def register(subparsers): + split = subparsers.add_parser("split") + split.add_argument("--input", required=True) + split.add_argument("--start-date") + split.add_argument("--end-date") + split.add_argument("--num-splits", type=int, default=3) + split.set_defaults(handler=execute) + common = subparsers.add_parser("common") + common.add_argument("--input", nargs="+", required=True) + common.add_argument("--keys", required=True) + common.add_argument("--labels") + common.add_argument("--output", required=True) + common.set_defaults(handler=execute) + charts = subparsers.add_parser("charts") + charts.add_argument("--input", nargs="+", required=True) + charts.add_argument("--output", required=True) + charts.set_defaults(handler=execute) + + +def execute(args): + if args.command == "split": + from gpu_backtest.workflows.splits import _read_input, split_csv_by_date_range + + data, column = _read_input(args.input, "Time") + split_csv_by_date_range( + args.input, + args.start_date or data[column].min().strftime("%d-%m-%Y"), + args.end_date or data[column].max().strftime("%d-%m-%Y"), + args.num_splits, + column, + ) + elif args.command == "common": + from gpu_backtest.workflows.common import process_data + + process_data( + { + "input_files": args.input, + "key_columns": [p.strip() for p in args.keys.split(",") if p.strip()], + "file_labels": args.labels.split(",") if args.labels else None, + "output_file": args.output, + } + ) + else: + from gpu_backtest.workflows.charts import write_charts + + print(write_charts(args.input, args.output)) diff --git a/src/gpu_backtest/cli/backtest.py b/src/gpu_backtest/cli/backtest.py new file mode 100644 index 0000000..7019e40 --- /dev/null +++ b/src/gpu_backtest/cli/backtest.py @@ -0,0 +1,42 @@ +"""Run commands: the GPU engine and generic split/common pipeline.""" + +from .options import job_options + + +def register(subparsers): + for name in ("run", "pipeline"): + command = subparsers.add_parser(name) + command.add_argument("--config") + command.add_argument("--strategy", help="Installed strategy module name") + command.add_argument("--input") + command.add_argument("--buy", type=float) + command.add_argument("--sell", type=float) + command.add_argument("--top-n", dest="top_n", type=int) + command.add_argument("--threads-per-block", dest="threads_per_block", type=int) + command.add_argument("--expected-interval", dest="expected_interval") + command.set_defaults(handler=execute) + if name == "run": + command.add_argument("--out-prefix", required=True) + else: + command.add_argument("--output-dir", required=True) + + +def execute(args): + strategy, source, options, config = job_options(args) + if args.command == "run": + from gpu_backtest.core.engine import run + + run(strategy, source, args.out_prefix, **options) + else: + from gpu_backtest.workflows.pipeline import run_pipeline + + run_pipeline( + strategy, + source, + args.output_dir, + num_splits=config.get("num_splits", 3), + start_date=config.get("start_date"), + end_date=config.get("end_date"), + ranges=config.get("ranges"), + **options, + ) diff --git a/src/gpu_backtest/cli/options.py b/src/gpu_backtest/cli/options.py new file mode 100644 index 0000000..6d2ac86 --- /dev/null +++ b/src/gpu_backtest/cli/options.py @@ -0,0 +1,33 @@ +"""Explicit config/CLI inputs, with relative config data paths.""" + +import json +from pathlib import Path + +from gpu_backtest import resolve_strategy_name + + +def job_options(args): + config = {} + if args.config: + config_path = Path(args.config).resolve() + config = json.loads(config_path.read_text()) + if not isinstance(config, dict): + raise ValueError("Config must be a JSON object") + if config.get("input"): + config["input"] = str(config_path.parent / config["input"]) + strategy = args.strategy if args.strategy is not None else config.get("strategy") + source = args.input if args.input is not None else config.get("input") + if not strategy or not source: + raise ValueError("Specify strategy and input explicitly, either in config or CLI flags") + options = {} + for name, default in ( + ("buy", 0.0015), + ("sell", 0.0015), + ("top_n", 100), + ("threads_per_block", 128), + ("expected_interval", None), + ): + value = getattr(args, name) + options[name] = config.get(name, default) if value is None else value + options.update(entry_dims=config.get("entry_dims"), exit_dims=config.get("exit_dims")) + return (resolve_strategy_name(strategy), source, options, config) diff --git a/src/gpu_backtest/cli/tools.py b/src/gpu_backtest/cli/tools.py new file mode 100644 index 0000000..2c9c9d4 --- /dev/null +++ b/src/gpu_backtest/cli/tools.py @@ -0,0 +1,73 @@ +"""CLI adapters for optional tools; they do not implement engine computation.""" + + +def register(subparsers): + check = subparsers.add_parser("gpu-check", help="Real-GPU numerical smoke checks") + check.add_argument("--output") + check.set_defaults(handler=execute) + benchmark = subparsers.add_parser( + "benchmark", help="Public RSI CPU/GPU performance measurement" + ) + benchmark.add_argument("--output", required=True) + benchmark.add_argument("--bars", type=int, default=1024) + benchmark.add_argument("--cpu-threads", type=int, default=8) + benchmark.add_argument("--repeats", type=int, default=3) + benchmark.add_argument("--billion", action="store_true") + benchmark.add_argument("--billion-cpu", action="store_true") + benchmark.set_defaults(handler=execute) + from gpu_backtest_tools.runpod.launcher import DEFAULT_IMAGE + + runpod = subparsers.add_parser("runpod", help="Run a job on a leased GPU and clean it up") + runpod.add_argument("--config") + runpod.add_argument("--output-dir", required=True) + runpod.add_argument("--ssh-key") + runpod.add_argument("--plugin-dir") + runpod.add_argument( + "--mode", choices=("run", "pipeline", "check", "benchmark"), default="pipeline" + ) + runpod.add_argument("--image", default=DEFAULT_IMAGE) + runpod.add_argument("--gpu", default="NVIDIA GeForce RTX 4090") + runpod.add_argument("--cloud", choices=("SECURE", "COMMUNITY"), default="SECURE") + runpod.add_argument("--max-seconds", type=int, default=1800) + runpod.add_argument("--max-hourly-rate", type=float, default=1.0) + runpod.add_argument("--keep-pod", action="store_true") + runpod.add_argument("--charts", action="store_true") + runpod.add_argument("--dry-run", action="store_true") + runpod.set_defaults(handler=execute) + + +def execute(args): + if args.command == "gpu-check": + from gpu_backtest_tools.checks.gpu import check_gpu + + check_gpu(args.output) + elif args.command == "benchmark": + from gpu_backtest_tools.benchmarks.runner import benchmark + + benchmark( + args.output, + bars=args.bars, + cpu_threads=args.cpu_threads, + repeats=args.repeats, + billion=args.billion, + billion_cpu=args.billion_cpu, + ) + else: + from gpu_backtest_tools.runpod.launcher import launch, terminate_on_signal + + with terminate_on_signal(): + launch( + config_path=args.config, + output_dir=args.output_dir, + ssh_key=args.ssh_key, + mode=args.mode, + plugin_dir=args.plugin_dir, + image=args.image, + gpu=args.gpu, + cloud=args.cloud, + max_seconds=args.max_seconds, + max_hourly_rate=args.max_hourly_rate, + keep_pod=args.keep_pod, + charts=args.charts, + dry_run=args.dry_run, + ) diff --git a/src/gpu_backtest/core/__init__.py b/src/gpu_backtest/core/__init__.py new file mode 100644 index 0000000..11a9d5b --- /dev/null +++ b/src/gpu_backtest/core/__init__.py @@ -0,0 +1 @@ +"""Strategy-agnostic GPU computation and artifact primitives.""" diff --git a/src/gpu_backtest/core/data.py b/src/gpu_backtest/core/data.py new file mode 100644 index 0000000..85aee02 --- /dev/null +++ b/src/gpu_backtest/core/data.py @@ -0,0 +1,55 @@ +"""Load and validate market input without silently repairing it.""" + +import numpy as np +import pandas as pd + +REQUIRED_MARKET_COLUMNS = ("Time", "Open", "High", "Low", "Close", "Volume") + + +def validate_market_data(input_csv, expected_interval=None): + """Load and validate an OHLCV file without repairing invalid input.""" + data = pd.read_csv(input_csv) + missing = [column for column in REQUIRED_MARKET_COLUMNS if column not in data.columns] + if missing: + raise ValueError(f"{input_csv}: missing required columns: {missing}") + if len(data) < 2: + raise ValueError(f"{input_csv}: at least two bars are required") + try: + data["Time"] = pd.to_datetime(data["Time"], errors="raise") + except Exception as exc: + raise ValueError(f"{input_csv}: Time contains invalid timestamps") from exc + if data["Time"].isna().any(): + raise ValueError(f"{input_csv}: Time contains missing timestamps") + if data["Time"].duplicated().any(): + raise ValueError(f"{input_csv}: Time contains duplicate timestamps") + if not data["Time"].is_monotonic_increasing: + raise ValueError(f"{input_csv}: Time must be strictly increasing") + numeric = data[list(REQUIRED_MARKET_COLUMNS[1:])].apply(pd.to_numeric, errors="coerce") + values = numeric.to_numpy(dtype=np.float64) + if not np.isfinite(values).all(): + raise ValueError(f"{input_csv}: OHLCV values must all be finite numbers") + data.loc[:, list(REQUIRED_MARKET_COLUMNS[1:])] = numeric + prices = numeric[["Open", "High", "Low", "Close"]] + if (prices <= 0).any().any(): + raise ValueError(f"{input_csv}: OHLC prices must be positive") + if (numeric["Volume"] < 0).any(): + raise ValueError(f"{input_csv}: Volume must be non-negative") + if (numeric["High"] < prices[["Open", "Low", "Close"]].max(axis=1)).any(): + raise ValueError(f"{input_csv}: High is below another OHLC price") + if (numeric["Low"] > prices[["Open", "High", "Close"]].min(axis=1)).any(): + raise ValueError(f"{input_csv}: Low is above another OHLC price") + if expected_interval: + try: + interval = pd.to_timedelta(expected_interval) + except Exception as exc: + raise ValueError(f"Invalid expected interval: {expected_interval}") from exc + if interval <= pd.Timedelta(0): + raise ValueError("Expected interval must be positive") + deltas = data["Time"].diff().iloc[1:] + bad = deltas != interval + if bad.any(): + first_bad = int(np.flatnonzero(bad.to_numpy())[0]) + 1 + raise ValueError( + f"{input_csv}: bar interval at row {first_bad} is {deltas.iloc[first_bad - 1]}, expected {interval}" + ) + return data diff --git a/src/gpu_backtest/core/engine.py b/src/gpu_backtest/core/engine.py new file mode 100644 index 0000000..d69f387 --- /dev/null +++ b/src/gpu_backtest/core/engine.py @@ -0,0 +1,138 @@ +"""Public GPU runner; orchestration delegates to the focused core modules.""" + +import time + +import numpy as np +from numba import cuda + +from .data import validate_market_data +from .grid import dim_arrays, space_count +from .indicators import prepare_tables +from .kernels import build_kernels +from .output import validate_reduction_outputs, write_top_csv +from .strategy import MAX_DIMS, load_strategy, validate_strategy + + +def run( + strategy_name, + input_csv, + out_prefix, + buy=0.0015, + sell=0.0015, + top_n=10000, + threads_per_block=128, + entry_dims=None, + exit_dims=None, + expected_interval=None, + verbose=True, +): + strategy = load_strategy(strategy_name) + e_dims = strategy.ENTRY_DIMS if entry_dims is None else entry_dims + x_dims = strategy.EXIT_DIMS if exit_dims is None else exit_dims + if not e_dims or not x_dims or len(e_dims) > MAX_DIMS or (len(x_dims) > MAX_DIMS): + raise ValueError(f"Entry and exit dimensions must each contain 1 to {MAX_DIMS} dimensions") + if not np.isfinite([buy, sell]).all() or buy < 0 or sell < 0: + raise ValueError("Commissions must be finite and non-negative") + if isinstance(top_n, bool) or not isinstance(top_n, (int, np.integer)) or top_n < 1: + raise ValueError("top_n must be a positive integer") + if ( + isinstance(threads_per_block, bool) + or not isinstance(threads_per_block, (int, np.integer)) + or (not 1 <= threads_per_block <= 1024) + ): + raise ValueError("threads_per_block must be an integer between 1 and 1024") + validate_strategy(strategy, e_dims, x_dims) + entry_count, exit_count = (space_count(e_dims), space_count(x_dims)) + if out_prefix and (entry_count < 2 or exit_count < 2): + raise ValueError( + "Effect-size isolation requires at least two entry and two exit combinations" + ) + if verbose: + print( + f"[{strategy.NAME}] entry={entry_count} × exit={exit_count} = {entry_count * exit_count:,} combos" + ) + data = validate_market_data(input_csv, expected_interval=expected_interval) + close = data["Close"].values.astype(np.float32) + open_ = data["Open"].values.astype(np.float32) + high = data["High"].values.astype(np.float32) + low = data["Low"].values.astype(np.float32) + vol = data["Volume"].values.astype(np.float32) + for name, values in ( + ("Close", close), + ("Open", open_), + ("High", high), + ("Low", low), + ("Volume", vol), + ): + if not np.isfinite(values).all(): + raise ValueError(f"{input_csv}: {name} cannot be represented as finite float32 values") + num_bars = len(data) + + class _S: + ENTRY_DIMS, EXIT_DIMS, TABLES = (e_dims, x_dims, strategy.TABLES) + + tabs, rms = prepare_tables(_S, close, open_, high, low, vol) + d = [cuda.to_device(a) for a in (close, open_, high, low, vol)] + dt = [cuda.to_device(t) for t in tabs] + dr = [cuda.to_device(r) for r in rms] + e_lo, e_st, e_ct, n_e = dim_arrays(e_dims) + x_lo, x_st, x_ct, n_x = dim_arrays(x_dims) + de = [cuda.to_device(a) for a in (e_lo, e_st, e_ct)] + dx = [cuda.to_device(a) for a in (x_lo, x_st, x_ct)] + es = cuda.device_array(entry_count, np.float64) + esq = cuda.device_array(entry_count, np.float64) + xs = cuda.device_array(exit_count, np.float64) + xsq = cuda.device_array(exit_count, np.float64) + entry_k, exit_k = build_kernels(strategy) + args = ( + *d, + *dt, + *dr, + *de, + n_e, + *dx, + n_x, + entry_count, + exit_count, + num_bars, + np.float64(buy), + np.float64(sell), + ) + t0 = time.time() + eb = (entry_count + threads_per_block - 1) // threads_per_block + entry_k[eb, threads_per_block](*args, es, esq) + cuda.synchronize() + if verbose: + print(f"Pass1 (entry reduce) {time.time() - t0:.2f}s") + t1 = time.time() + xb = (exit_count + threads_per_block - 1) // threads_per_block + exit_k[xb, threads_per_block](*args, xs, xsq) + cuda.synchronize() + if verbose: + print(f"Pass2 (exit reduce) {time.time() - t1:.2f}s") + es, esq, xs, xsq = ( + es.copy_to_host(), + esq.copy_to_host(), + xs.copy_to_host(), + xsq.copy_to_host(), + ) + reduction_checks = validate_reduction_outputs(es, esq, xs, xsq) + if verbose: + print(f"grand total check: diff={reduction_checks['sum']['difference']:.2e}") + out = {"entry_sum": es, "entry_sumsq": esq, "exit_sum": xs, "exit_sumsq": xsq} + if out_prefix: + write_top_csv( + out_prefix, + strategy.NAME, + input_csv, + e_dims, + x_dims, + es, + esq, + xs, + xsq, + top_n, + reduction_checks, + verbose, + ) + return out diff --git a/src/gpu_backtest/core/grid.py b/src/gpu_backtest/core/grid.py new file mode 100644 index 0000000..c3213bc --- /dev/null +++ b/src/gpu_backtest/core/grid.py @@ -0,0 +1,51 @@ +"""Parameter counts, decoding, and indicator-window selection.""" + +import numpy as np + +from .strategy import MAX_DIMS + + +def dim_count(d): + name, lo, hi, step, is_float = d + return int(round((hi - lo) / step) + 1 if is_float else (hi - lo) // step + 1) + + +def space_count(dims): + n = 1 + for d in dims: + n *= dim_count(d) + return n + + +def dim_arrays(dims): + """Return padded lower bounds, steps, counts, and dimension count.""" + lo = np.zeros(MAX_DIMS, np.float64) + st = np.ones(MAX_DIMS, np.float64) + ct = np.ones(MAX_DIMS, np.int64) + for j, d in enumerate(dims): + lo[j] = float(d[1]) + st[j] = float(d[3]) + ct[j] = dim_count(d) + return (lo, st, ct, len(dims)) + + +def decode_values(dims, idx): + """Decode a flat index; the last dimension varies fastest.""" + cnts = [dim_count(d) for d in dims] + vals = [0.0] * len(dims) + rem = idx + for j in range(len(dims) - 1, -1, -1): + i = rem % cnts[j] + rem //= cnts[j] + v = dims[j][1] + i * dims[j][3] + vals[j] = float(v) if dims[j][4] else int(v) + return vals + + +def int_dim_values(dims, names): + out = set() + for d in dims: + if d[0] in names: + assert not d[4], f"Table window dimension {d[0]} must contain integers" + out |= {int(d[1] + i * d[3]) for i in range(dim_count(d))} + return sorted(out) diff --git a/src/gpu_backtest/tables.py b/src/gpu_backtest/core/indicators.py similarity index 75% rename from src/gpu_backtest/tables.py rename to src/gpu_backtest/core/indicators.py index 368ab7b..f446cd8 100644 --- a/src/gpu_backtest/tables.py +++ b/src/gpu_backtest/core/indicators.py @@ -3,6 +3,9 @@ import numpy as np import pandas as pd +from .grid import int_dim_values +from .strategy import MAX_TABLES + def _series(kind, source, df): if kind == "atr": @@ -58,3 +61,18 @@ def build_table(kind, source, windows, close, open_, high, low, volume): for i, w in enumerate(windows): row_map[w] = i return (tab, row_map) + + +def prepare_tables(strategy, close, open_, high, low, vol): + all_dims = list(strategy.ENTRY_DIMS) + list(strategy.EXIT_DIMS) + tabs, rms = ([], []) + for kind, source, dim_names in strategy.TABLES: + windows = int_dim_values(all_dims, dim_names) + assert windows, f"Table {kind} has no integer window values for {dim_names}" + t, rm = build_table(kind, source, windows, close, open_, high, low, vol) + tabs.append(t) + rms.append(rm) + while len(tabs) < MAX_TABLES: + tabs.append(np.zeros((1, 1), np.float32)) + rms.append(np.zeros(1, np.int64)) + return (tabs, rms) diff --git a/src/gpu_backtest/core/kernels.py b/src/gpu_backtest/core/kernels.py new file mode 100644 index 0000000..380e280 --- /dev/null +++ b/src/gpu_backtest/core/kernels.py @@ -0,0 +1,151 @@ +"""GPU kernels: two deterministic reductions over strategy returns.""" + +from numba import cuda, float32, float64 + + +def build_kernels(strategy): + algo = strategy.make_device_fn(cuda) + + @cuda.jit(device=True, inline=True) + def _decode(idx, lo, st, ct, nd, out): + rem = idx + for j in range(nd - 1, -1, -1): + i = rem % ct[j] + rem //= ct[j] + out[j] = lo[j] + i * st[j] + + @cuda.jit + def entry_reduce( + close, + open_, + high, + low, + vol, + t0, + t1, + t2, + t3, + rm0, + rm1, + rm2, + rm3, + e_lo, + e_st, + e_ct, + n_e, + x_lo, + x_st, + x_ct, + n_x, + entry_count, + exit_count, + num_bars, + buy, + sell, + out_sum, + out_sumsq, + ): + e = cuda.grid(1) + if e >= entry_count: + return + ep = cuda.local.array(4, float64) + xp = cuda.local.array(4, float64) + _decode(e, e_lo, e_st, e_ct, n_e, ep) + s = float64(0.0) + sq = float64(0.0) + for x in range(exit_count): + _decode(x, x_lo, x_st, x_ct, n_x, xp) + tr64 = algo( + close, + open_, + high, + low, + vol, + t0, + t1, + t2, + t3, + rm0, + rm1, + rm2, + rm3, + ep, + xp, + num_bars, + buy, + sell, + ) + tr = float32(tr64) + s += float64(tr) + sq += float64(float32(tr * tr)) + out_sum[e] = s + out_sumsq[e] = sq + + @cuda.jit + def exit_reduce( + close, + open_, + high, + low, + vol, + t0, + t1, + t2, + t3, + rm0, + rm1, + rm2, + rm3, + e_lo, + e_st, + e_ct, + n_e, + x_lo, + x_st, + x_ct, + n_x, + entry_count, + exit_count, + num_bars, + buy, + sell, + out_sum, + out_sumsq, + ): + x = cuda.grid(1) + if x >= exit_count: + return + ep = cuda.local.array(4, float64) + xp = cuda.local.array(4, float64) + _decode(x, x_lo, x_st, x_ct, n_x, xp) + s = float64(0.0) + sq = float64(0.0) + for e in range(entry_count): + _decode(e, e_lo, e_st, e_ct, n_e, ep) + tr64 = algo( + close, + open_, + high, + low, + vol, + t0, + t1, + t2, + t3, + rm0, + rm1, + rm2, + rm3, + ep, + xp, + num_bars, + buy, + sell, + ) + tr = float32(tr64) + s += float64(tr) + sq += float64(float32(tr * tr)) + out_sum[x] = s + out_sumsq[x] = sq + + return (entry_reduce, exit_reduce) diff --git a/src/gpu_backtest/core/output.py b/src/gpu_backtest/core/output.py new file mode 100644 index 0000000..efdedfe --- /dev/null +++ b/src/gpu_backtest/core/output.py @@ -0,0 +1,125 @@ +"""Validate grouped outputs and write ranked CSVs and their manifest.""" + +import json +import os +from pathlib import Path + +import numpy as np +import pandas as pd + +from . import statistics as cs +from .grid import decode_values + +ROUND_TRIP_FLOAT_FORMAT = "%.17g" + + +def validate_reduction_outputs(entry_sum, entry_sumsq, exit_sum, exit_sumsq): + """Check that both reduction passes have matching grand totals.""" + arrays = { + "entry_sum": np.asarray(entry_sum, dtype=np.float64), + "entry_sumsq": np.asarray(entry_sumsq, dtype=np.float64), + "exit_sum": np.asarray(exit_sum, dtype=np.float64), + "exit_sumsq": np.asarray(exit_sumsq, dtype=np.float64), + } + for name, values in arrays.items(): + if values.size == 0 or not np.isfinite(values).all(): + raise RuntimeError(f"{name} is empty or contains non-finite values") + if (arrays["entry_sumsq"] < 0).any() or (arrays["exit_sumsq"] < 0).any(): + raise RuntimeError("Reduction sum-of-squares values must be non-negative") + checks = {} + for name, left, right in ( + ( + "sum", + arrays["entry_sum"].sum(dtype=np.float64), + arrays["exit_sum"].sum(dtype=np.float64), + ), + ( + "sumsq", + arrays["entry_sumsq"].sum(dtype=np.float64), + arrays["exit_sumsq"].sum(dtype=np.float64), + ), + ): + difference = abs(left - right) + tolerance = 1e-06 + 1e-12 * max(abs(left), abs(right), 1.0) + if difference > tolerance: + raise RuntimeError( + f"Entry/exit grand {name} mismatch: diff={difference:.17g}, tolerance={tolerance:.17g}" + ) + checks[name] = {"entry": float(left), "exit": float(right), "difference": float(difference)} + return checks + + +def _format_parameter(value): + return np.format_float_positional(float(value), trim="-") + + +def write_top_csv( + prefix, + strategy_name, + input_csv, + e_dims, + x_dims, + es, + esq, + xs, + xsq, + top_n, + reduction_checks, + verbose, +): + output_records = {} + for side, dims, s, sq, n1 in ( + ("entry", e_dims, es, esq, len(xs)), + ("exit", x_dims, xs, xsq, len(es)), + ): + st = cs.stats_from_groups(s, sq, n1, s.sum(), sq.sum(), len(s) * n1) + cols = { + d[0]: np.array([decode_values(dims, i)[j] for i in range(len(s))]) + for j, d in enumerate(dims) + } + df = pd.DataFrame({**cols, **st}) + if not np.isfinite(df["effect_size"].to_numpy(dtype=np.float64)).all(): + raise RuntimeError(f"{side} effect_size contains non-finite values") + key_columns = [d[0] for d in dims] + df = df.sort_values( + ["effect_size", *key_columns], + ascending=[False, *[True] * len(key_columns)], + kind="mergesort", + ).head(min(top_n, len(s))) + serialized = df.copy() + for dim in dims: + if dim[4]: + serialized[dim[0]] = serialized[dim[0]].map(_format_parameter) + path = f"{prefix}_top_{side}.csv" + Path(path).parent.mkdir(parents=True, exist_ok=True) + temp_path = f"{path}.tmp" + serialized.to_csv(temp_path, index=False, float_format=ROUND_TRIP_FLOAT_FORMAT) + os.replace(temp_path, path) + output_records[side] = { + "file": Path(path).name, + "rows": int(len(df)), + "key_columns": key_columns, + "columns": list(serialized.columns), + "sort": ["effect_size descending", "parameter columns ascending"], + } + if verbose: + print(f" -> {path} ({len(df)} rows)") + manifest = { + "schema_version": 1, + "engine": "gpu_backtest.engine", + "strategy": strategy_name, + "source_file": Path(input_csv).name, + "entry_count": int(len(es)), + "exit_count": int(len(xs)), + "pair_count": int(len(es) * len(xs)), + "top_n": int(top_n), + "float_serialization": "IEEE-754 float64 round-trip (17 significant digits)", + "reduction_checks": reduction_checks, + "outputs": output_records, + } + manifest_path = f"{prefix}_top_manifest.json" + temp_manifest = f"{manifest_path}.tmp" + Path(temp_manifest).write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8") + os.replace(temp_manifest, manifest_path) + if verbose: + print(f" -> {manifest_path}") diff --git a/src/gpu_backtest/statistics.py b/src/gpu_backtest/core/statistics.py similarity index 100% rename from src/gpu_backtest/statistics.py rename to src/gpu_backtest/core/statistics.py diff --git a/src/gpu_backtest/strategy.py b/src/gpu_backtest/core/strategy.py similarity index 94% rename from src/gpu_backtest/strategy.py rename to src/gpu_backtest/core/strategy.py index fd3565a..4becefd 100644 --- a/src/gpu_backtest/strategy.py +++ b/src/gpu_backtest/core/strategy.py @@ -14,11 +14,11 @@ def load_strategy(value): - """A short name selects a builtin; a dotted name imports an external module.""" + """Import a strategy module or accept an already loaded strategy object.""" if isinstance(value, str): if not value or not all(part.isidentifier() for part in value.split(".")): - raise ValueError("Strategy must be a builtin name or a dotted Python module name") - module_name = value if "." in value else f"gpu_backtest.strategies.{value}" + raise ValueError("Strategy must be a Python module name") + module_name = value value = importlib.import_module(module_name) for attribute in ("NAME", "ENTRY_DIMS", "EXIT_DIMS", "TABLES", "make_device_fn"): if not hasattr(value, attribute): diff --git a/src/gpu_backtest/engine.py b/src/gpu_backtest/engine.py deleted file mode 100644 index 0e1f020..0000000 --- a/src/gpu_backtest/engine.py +++ /dev/null @@ -1,541 +0,0 @@ -"""Two-pass GPU parameter sweeps with deterministic grouped statistics.""" - -import json -import os -import time -from pathlib import Path - -import numpy as np -import pandas as pd -from numba import cuda, float32, float64 - -from . import statistics as cs -from .strategy import load_strategy, validate_strategy -from .tables import build_table - -MAX_TABLES = 4 -MAX_DIMS = 4 -INITIAL_CAPITAL = 10000.0 -REQUIRED_MARKET_COLUMNS = ("Time", "Open", "High", "Low", "Close", "Volume") -ROUND_TRIP_FLOAT_FORMAT = "%.17g" - - -def dim_count(d): - name, lo, hi, step, is_float = d - return int(round((hi - lo) / step) + 1 if is_float else (hi - lo) // step + 1) - - -def space_count(dims): - n = 1 - for d in dims: - n *= dim_count(d) - return n - - -def validate_market_data(input_csv, expected_interval=None): - """Load and validate an OHLCV file without repairing invalid input.""" - data = pd.read_csv(input_csv) - missing = [column for column in REQUIRED_MARKET_COLUMNS if column not in data.columns] - if missing: - raise ValueError(f"{input_csv}: missing required columns: {missing}") - if len(data) < 2: - raise ValueError(f"{input_csv}: at least two bars are required") - try: - data["Time"] = pd.to_datetime(data["Time"], errors="raise") - except Exception as exc: - raise ValueError(f"{input_csv}: Time contains invalid timestamps") from exc - if data["Time"].isna().any(): - raise ValueError(f"{input_csv}: Time contains missing timestamps") - if data["Time"].duplicated().any(): - raise ValueError(f"{input_csv}: Time contains duplicate timestamps") - if not data["Time"].is_monotonic_increasing: - raise ValueError(f"{input_csv}: Time must be strictly increasing") - numeric = data[list(REQUIRED_MARKET_COLUMNS[1:])].apply(pd.to_numeric, errors="coerce") - values = numeric.to_numpy(dtype=np.float64) - if not np.isfinite(values).all(): - raise ValueError(f"{input_csv}: OHLCV values must all be finite numbers") - data.loc[:, list(REQUIRED_MARKET_COLUMNS[1:])] = numeric - prices = numeric[["Open", "High", "Low", "Close"]] - if (prices <= 0).any().any(): - raise ValueError(f"{input_csv}: OHLC prices must be positive") - if (numeric["Volume"] < 0).any(): - raise ValueError(f"{input_csv}: Volume must be non-negative") - if (numeric["High"] < prices[["Open", "Low", "Close"]].max(axis=1)).any(): - raise ValueError(f"{input_csv}: High is below another OHLC price") - if (numeric["Low"] > prices[["Open", "High", "Close"]].min(axis=1)).any(): - raise ValueError(f"{input_csv}: Low is above another OHLC price") - if expected_interval: - try: - interval = pd.to_timedelta(expected_interval) - except Exception as exc: - raise ValueError(f"Invalid expected interval: {expected_interval}") from exc - if interval <= pd.Timedelta(0): - raise ValueError("Expected interval must be positive") - deltas = data["Time"].diff().iloc[1:] - bad = deltas != interval - if bad.any(): - first_bad = int(np.flatnonzero(bad.to_numpy())[0]) + 1 - raise ValueError( - f"{input_csv}: bar interval at row {first_bad} is {deltas.iloc[first_bad - 1]}, expected {interval}" - ) - return data - - -def validate_reduction_outputs(entry_sum, entry_sumsq, exit_sum, exit_sumsq): - """Check that both reduction passes have matching grand totals.""" - arrays = { - "entry_sum": np.asarray(entry_sum, dtype=np.float64), - "entry_sumsq": np.asarray(entry_sumsq, dtype=np.float64), - "exit_sum": np.asarray(exit_sum, dtype=np.float64), - "exit_sumsq": np.asarray(exit_sumsq, dtype=np.float64), - } - for name, values in arrays.items(): - if values.size == 0 or not np.isfinite(values).all(): - raise RuntimeError(f"{name} is empty or contains non-finite values") - if (arrays["entry_sumsq"] < 0).any() or (arrays["exit_sumsq"] < 0).any(): - raise RuntimeError("Reduction sum-of-squares values must be non-negative") - checks = {} - for name, left, right in ( - ( - "sum", - arrays["entry_sum"].sum(dtype=np.float64), - arrays["exit_sum"].sum(dtype=np.float64), - ), - ( - "sumsq", - arrays["entry_sumsq"].sum(dtype=np.float64), - arrays["exit_sumsq"].sum(dtype=np.float64), - ), - ): - difference = abs(left - right) - tolerance = 1e-06 + 1e-12 * max(abs(left), abs(right), 1.0) - if difference > tolerance: - raise RuntimeError( - f"Entry/exit grand {name} mismatch: diff={difference:.17g}, tolerance={tolerance:.17g}" - ) - checks[name] = {"entry": float(left), "exit": float(right), "difference": float(difference)} - return checks - - -def dim_arrays(dims): - """Return padded lower bounds, steps, counts, and dimension count.""" - lo = np.zeros(MAX_DIMS, np.float64) - st = np.ones(MAX_DIMS, np.float64) - ct = np.ones(MAX_DIMS, np.int64) - for j, d in enumerate(dims): - lo[j] = float(d[1]) - st[j] = float(d[3]) - ct[j] = dim_count(d) - return (lo, st, ct, len(dims)) - - -def decode_values(dims, idx): - """Decode a flat index; the last dimension varies fastest.""" - cnts = [dim_count(d) for d in dims] - vals = [0.0] * len(dims) - rem = idx - for j in range(len(dims) - 1, -1, -1): - i = rem % cnts[j] - rem //= cnts[j] - v = dims[j][1] + i * dims[j][3] - vals[j] = float(v) if dims[j][4] else int(v) - return vals - - -def int_dim_values(dims, names): - out = set() - for d in dims: - if d[0] in names: - assert not d[4], f"Table window dimension {d[0]} must contain integers" - out |= {int(d[1] + i * d[3]) for i in range(dim_count(d))} - return sorted(out) - - -def prepare_tables(strategy, close, open_, high, low, vol): - all_dims = list(strategy.ENTRY_DIMS) + list(strategy.EXIT_DIMS) - tabs, rms = ([], []) - for kind, source, dim_names in strategy.TABLES: - windows = int_dim_values(all_dims, dim_names) - assert windows, f"Table {kind} has no integer window values for {dim_names}" - t, rm = build_table(kind, source, windows, close, open_, high, low, vol) - tabs.append(t) - rms.append(rm) - while len(tabs) < MAX_TABLES: - tabs.append(np.zeros((1, 1), np.float32)) - rms.append(np.zeros(1, np.int64)) - return (tabs, rms) - - -def build_kernels(strategy): - algo = strategy.make_device_fn(cuda) - - @cuda.jit(device=True, inline=True) - def _decode(idx, lo, st, ct, nd, out): - rem = idx - for j in range(nd - 1, -1, -1): - i = rem % ct[j] - rem //= ct[j] - out[j] = lo[j] + i * st[j] - - @cuda.jit - def entry_reduce( - close, - open_, - high, - low, - vol, - t0, - t1, - t2, - t3, - rm0, - rm1, - rm2, - rm3, - e_lo, - e_st, - e_ct, - n_e, - x_lo, - x_st, - x_ct, - n_x, - entry_count, - exit_count, - num_bars, - buy, - sell, - out_sum, - out_sumsq, - ): - e = cuda.grid(1) - if e >= entry_count: - return - ep = cuda.local.array(4, float64) - xp = cuda.local.array(4, float64) - _decode(e, e_lo, e_st, e_ct, n_e, ep) - s = float64(0.0) - sq = float64(0.0) - for x in range(exit_count): - _decode(x, x_lo, x_st, x_ct, n_x, xp) - tr64 = algo( - close, - open_, - high, - low, - vol, - t0, - t1, - t2, - t3, - rm0, - rm1, - rm2, - rm3, - ep, - xp, - num_bars, - buy, - sell, - ) - tr = float32(tr64) - s += float64(tr) - sq += float64(float32(tr * tr)) - out_sum[e] = s - out_sumsq[e] = sq - - @cuda.jit - def exit_reduce( - close, - open_, - high, - low, - vol, - t0, - t1, - t2, - t3, - rm0, - rm1, - rm2, - rm3, - e_lo, - e_st, - e_ct, - n_e, - x_lo, - x_st, - x_ct, - n_x, - entry_count, - exit_count, - num_bars, - buy, - sell, - out_sum, - out_sumsq, - ): - x = cuda.grid(1) - if x >= exit_count: - return - ep = cuda.local.array(4, float64) - xp = cuda.local.array(4, float64) - _decode(x, x_lo, x_st, x_ct, n_x, xp) - s = float64(0.0) - sq = float64(0.0) - for e in range(entry_count): - _decode(e, e_lo, e_st, e_ct, n_e, ep) - tr64 = algo( - close, - open_, - high, - low, - vol, - t0, - t1, - t2, - t3, - rm0, - rm1, - rm2, - rm3, - ep, - xp, - num_bars, - buy, - sell, - ) - tr = float32(tr64) - s += float64(tr) - sq += float64(float32(tr * tr)) - out_sum[x] = s - out_sumsq[x] = sq - - return (entry_reduce, exit_reduce) - - -def run( - strategy_name, - input_csv, - out_prefix, - buy=0.0015, - sell=0.0015, - top_n=10000, - threads_per_block=128, - entry_dims=None, - exit_dims=None, - expected_interval=None, - verbose=True, -): - strategy = load_strategy(strategy_name) - e_dims = strategy.ENTRY_DIMS if entry_dims is None else entry_dims - x_dims = strategy.EXIT_DIMS if exit_dims is None else exit_dims - if not e_dims or not x_dims or len(e_dims) > MAX_DIMS or (len(x_dims) > MAX_DIMS): - raise ValueError(f"Entry and exit dimensions must each contain 1 to {MAX_DIMS} dimensions") - if not np.isfinite([buy, sell]).all() or buy < 0 or sell < 0: - raise ValueError("Commissions must be finite and non-negative") - if isinstance(top_n, bool) or not isinstance(top_n, (int, np.integer)) or top_n < 1: - raise ValueError("top_n must be a positive integer") - if ( - isinstance(threads_per_block, bool) - or not isinstance(threads_per_block, (int, np.integer)) - or not 1 <= threads_per_block <= 1024 - ): - raise ValueError("threads_per_block must be an integer between 1 and 1024") - validate_strategy(strategy, e_dims, x_dims) - entry_count, exit_count = (space_count(e_dims), space_count(x_dims)) - if out_prefix and (entry_count < 2 or exit_count < 2): - raise ValueError( - "Effect-size isolation requires at least two entry and two exit combinations" - ) - if verbose: - print( - f"[{strategy.NAME}] entry={entry_count} × exit={exit_count} = {entry_count * exit_count:,} combos" - ) - data = validate_market_data(input_csv, expected_interval=expected_interval) - close = data["Close"].values.astype(np.float32) - open_ = data["Open"].values.astype(np.float32) - high = data["High"].values.astype(np.float32) - low = data["Low"].values.astype(np.float32) - vol = data["Volume"].values.astype(np.float32) - for name, values in ( - ("Close", close), - ("Open", open_), - ("High", high), - ("Low", low), - ("Volume", vol), - ): - if not np.isfinite(values).all(): - raise ValueError(f"{input_csv}: {name} cannot be represented as finite float32 values") - num_bars = len(data) - - class _S: - ENTRY_DIMS, EXIT_DIMS, TABLES = (e_dims, x_dims, strategy.TABLES) - - tabs, rms = prepare_tables(_S, close, open_, high, low, vol) - d = [cuda.to_device(a) for a in (close, open_, high, low, vol)] - dt = [cuda.to_device(t) for t in tabs] - dr = [cuda.to_device(r) for r in rms] - e_lo, e_st, e_ct, n_e = dim_arrays(e_dims) - x_lo, x_st, x_ct, n_x = dim_arrays(x_dims) - de = [cuda.to_device(a) for a in (e_lo, e_st, e_ct)] - dx = [cuda.to_device(a) for a in (x_lo, x_st, x_ct)] - es = cuda.device_array(entry_count, np.float64) - esq = cuda.device_array(entry_count, np.float64) - xs = cuda.device_array(exit_count, np.float64) - xsq = cuda.device_array(exit_count, np.float64) - entry_k, exit_k = build_kernels(strategy) - args = ( - *d, - *dt, - *dr, - *de, - n_e, - *dx, - n_x, - entry_count, - exit_count, - num_bars, - np.float64(buy), - np.float64(sell), - ) - t0 = time.time() - eb = (entry_count + threads_per_block - 1) // threads_per_block - entry_k[eb, threads_per_block](*args, es, esq) - cuda.synchronize() - if verbose: - print(f"Pass1 (entry reduce) {time.time() - t0:.2f}s") - t1 = time.time() - xb = (exit_count + threads_per_block - 1) // threads_per_block - exit_k[xb, threads_per_block](*args, xs, xsq) - cuda.synchronize() - if verbose: - print(f"Pass2 (exit reduce) {time.time() - t1:.2f}s") - es, esq, xs, xsq = ( - es.copy_to_host(), - esq.copy_to_host(), - xs.copy_to_host(), - xsq.copy_to_host(), - ) - reduction_checks = validate_reduction_outputs(es, esq, xs, xsq) - if verbose: - print(f"grand total check: diff={reduction_checks['sum']['difference']:.2e}") - out = {"entry_sum": es, "entry_sumsq": esq, "exit_sum": xs, "exit_sumsq": xsq} - if out_prefix: - _write_top_csv( - out_prefix, - strategy.NAME, - input_csv, - e_dims, - x_dims, - es, - esq, - xs, - xsq, - top_n, - reduction_checks, - verbose, - ) - return out - - -def _format_parameter(value): - return np.format_float_positional(float(value), trim="-") - - -def _write_top_csv( - prefix, - strategy_name, - input_csv, - e_dims, - x_dims, - es, - esq, - xs, - xsq, - top_n, - reduction_checks, - verbose, -): - output_records = {} - for side, dims, s, sq, n1 in ( - ("entry", e_dims, es, esq, len(xs)), - ("exit", x_dims, xs, xsq, len(es)), - ): - st = cs.stats_from_groups(s, sq, n1, s.sum(), sq.sum(), len(s) * n1) - cols = { - d[0]: np.array([decode_values(dims, i)[j] for i in range(len(s))]) - for j, d in enumerate(dims) - } - df = pd.DataFrame({**cols, **st}) - if not np.isfinite(df["effect_size"].to_numpy(dtype=np.float64)).all(): - raise RuntimeError(f"{side} effect_size contains non-finite values") - key_columns = [d[0] for d in dims] - df = df.sort_values( - ["effect_size", *key_columns], - ascending=[False, *[True] * len(key_columns)], - kind="mergesort", - ).head(min(top_n, len(s))) - serialized = df.copy() - for dim in dims: - if dim[4]: - serialized[dim[0]] = serialized[dim[0]].map(_format_parameter) - path = f"{prefix}_top_{side}.csv" - Path(path).parent.mkdir(parents=True, exist_ok=True) - temp_path = f"{path}.tmp" - serialized.to_csv(temp_path, index=False, float_format=ROUND_TRIP_FLOAT_FORMAT) - os.replace(temp_path, path) - output_records[side] = { - "file": Path(path).name, - "rows": int(len(df)), - "key_columns": key_columns, - "columns": list(serialized.columns), - "sort": ["effect_size descending", "parameter columns ascending"], - } - if verbose: - print(f" -> {path} ({len(df)} rows)") - manifest = { - "schema_version": 1, - "engine": "gpu_backtest.engine", - "strategy": strategy_name, - "source_file": Path(input_csv).name, - "entry_count": int(len(es)), - "exit_count": int(len(xs)), - "pair_count": int(len(es) * len(xs)), - "top_n": int(top_n), - "float_serialization": "IEEE-754 float64 round-trip (17 significant digits)", - "reduction_checks": reduction_checks, - "outputs": output_records, - } - manifest_path = f"{prefix}_top_manifest.json" - temp_manifest = f"{manifest_path}.tmp" - Path(temp_manifest).write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8") - os.replace(temp_manifest, manifest_path) - if verbose: - print(f" -> {manifest_path}") - - -def reference_group_sums(strategy, close, open_, high, low, vol, e_dims, x_dims, buy, sell): - """Compute small-grid grouped returns with a strategy CPU reference.""" - - strategy = validate_strategy(strategy, e_dims, x_dims) - if not callable(getattr(strategy, "reference", None)): - raise ValueError("CPU comparisons require a callable strategy.reference") - - class _S: - ENTRY_DIMS, EXIT_DIMS, TABLES = (e_dims, x_dims, strategy.TABLES) - - tabs, rms = prepare_tables(_S, close, open_, high, low, vol) - ec, xc = (space_count(e_dims), space_count(x_dims)) - es = np.zeros(ec) - xs = np.zeros(xc) - for e in range(ec): - ep = decode_values(e_dims, e) + [0.0] * (MAX_DIMS - len(e_dims)) - for x in range(xc): - xp = decode_values(x_dims, x) + [0.0] * (MAX_DIMS - len(x_dims)) - tr = np.float32( - strategy.reference(close, open_, high, low, vol, tabs, rms, ep, xp, buy, sell) - ) - es[e] += float(tr) - xs[x] += float(tr) - return (es, xs) diff --git a/src/gpu_backtest/runpod.py b/src/gpu_backtest/runpod.py deleted file mode 100644 index ebfef56..0000000 --- a/src/gpu_backtest/runpod.py +++ /dev/null @@ -1,615 +0,0 @@ -"""Run a selected config on one leased RunPod GPU and reclaim the pod afterwards.""" - -import contextlib -import hashlib -import importlib.metadata -import ipaddress -import json -import math -import os -import shlex -import shutil -import signal -import subprocess -import tarfile -import tempfile -import threading -import time -import tomllib -import urllib.error -import urllib.request -import uuid -from pathlib import Path - -from . import __version__ - -API_URL = "https://rest.runpod.io/v1" -DEFAULT_IMAGE = "runpod/pytorch:2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04" -ENGINE_FILES = ( - "__init__.py", - "__main__.py", - "benchmark.py", - "benchmark_cpu.py", - "charts.py", - "cli.py", - "common.py", - "engine.py", - "gpu_check.py", - "pipeline.py", - "runpod.py", - "splits.py", - "statistics.py", - "strategy.py", - "synthetic.py", - "tables.py", - "strategies/__init__.py", - "strategies/rsi_meanrev.py", -) -CONFIG_KEYS = { - "strategy", - "input", - "buy", - "sell", - "entry_dims", - "exit_dims", - "top_n", - "threads_per_block", - "expected_interval", - "num_splits", - "start_date", - "end_date", - "ranges", -} -IGNORED_PARTS = {".git", ".venv", "venv", "__pycache__", ".pytest_cache", ".ruff_cache"} - - -class RunPodError(RuntimeError): - pass - - -def load_api_key(): - key = os.environ.get("RUNPOD_API_KEY") - if not key: - path = Path.home() / ".runpod/config.toml" - if path.is_file(): - key = tomllib.loads(path.read_text()).get("apikey") - if not isinstance(key, str) or not key.strip(): - raise ValueError("Set RUNPOD_API_KEY or apikey in ~/.runpod/config.toml") - return key.strip() - - -class RunPodClient: - def __init__(self, key): - self.key = key - - def request(self, method, path, body=None): - request = urllib.request.Request( - API_URL + path, - data=json.dumps(body).encode() if body is not None else None, - headers={ - "Authorization": f"Bearer {self.key}", - "Content-Type": "application/json", - "Accept": "application/json", - "User-Agent": f"gpu-backtest-engine/{__version__}", - }, - method=method, - ) - try: - with urllib.request.urlopen(request, timeout=30) as response: - payload = response.read() - return json.loads(payload) if payload else None - except urllib.error.HTTPError as exc: - if method == "DELETE" and exc.code == 404: - return None - # Do not echo server bodies, auth headers, or credentials into logs. - raise RunPodError(f"RunPod {method} {path}: HTTP {exc.code}") from None - except (urllib.error.URLError, TimeoutError, json.JSONDecodeError) as exc: - raise RunPodError(f"RunPod {method} {path}: connection or response failure") from exc - - def create(self, body): - try: - result = self.request("POST", "/pods", body) - if ( - not isinstance(result, dict) - or not isinstance(result.get("id"), str) - or not result["id"].isascii() - or not result["id"].isalnum() - ): - raise RunPodError("Creation response did not contain a valid pod ID") - except RunPodError as exc: - if "HTTP 4" in str(exc): - raise - # POST is never blindly retried: its outcome may already be a paid pod. - try: - pods = self.request("GET", "/pods") - matches = [p for p in pods if p.get("name") == body["name"]] - except (RunPodError, TypeError): - matches = [] - if ( - len(matches) == 1 - and isinstance(matches[0].get("id"), str) - and matches[0]["id"].isascii() - and matches[0]["id"].isalnum() - ): - return matches[0] - raise RunPodError( - f"Creation outcome unknown; check RunPod for launch {body['name']}. " - "No second creation request was sent." - ) from exc - return result - - def get(self, pod_id): - return self.request("GET", "/pods/" + pod_id) - - def delete(self, pod_id): - for attempt in range(3): - try: - self.request("DELETE", "/pods/" + pod_id) - return - except RunPodError: - if attempt == 2: - raise - time.sleep(attempt + 1) - - -def endpoint(pod): - address = pod.get("publicIp") - port = (pod.get("portMappings") or {}).get("22") - if not address or not port: - return None - try: - address = str(ipaddress.ip_address(address)) - port = int(port) - if not 1 <= port <= 65535: - raise ValueError - except (ValueError, TypeError): - raise RunPodError("RunPod returned invalid public SSH connection details") from None - return address, port - - -def _remaining(deadline): - seconds = deadline - time.monotonic() - if seconds <= 0: - raise RunPodError("RunPod job exceeded its local time limit") - return seconds - - -class SSHTransport: - def __init__(self, address, port, key, output_dir): - self.base = [ - "ssh", - "-i", - str(key), - "-p", - str(port), - "-o", - "BatchMode=yes", - "-o", - "IdentitiesOnly=yes", - "-o", - "ForwardAgent=no", - "-o", - "ClearAllForwardings=yes", - "-o", - "StrictHostKeyChecking=accept-new", - "-o", - f"UserKnownHostsFile={json.dumps(str(Path(output_dir) / 'known_hosts'))}", - "-o", - "ConnectTimeout=10", - "-o", - "ServerAliveInterval=15", - "-o", - "ServerAliveCountMax=3", - f"root@{address}", - ] - self.environment = os.environ.copy() - self.environment.pop("RUNPOD_API_KEY", None) - - def ready(self): - result = subprocess.run( - [*self.base, "true"], - stdin=subprocess.DEVNULL, - capture_output=True, - timeout=15, - env=self.environment, - ) - return result.returncode == 0 - - def _run(self, command, *, stdin=None, stdout=None, log=None, timeout): - with contextlib.ExitStack() as stack: - input_file = ( - stack.enter_context(Path(stdin).open("rb")) if stdin else subprocess.DEVNULL - ) - output_file = stack.enter_context(Path(stdout).open("wb")) if stdout else None - log_file = stack.enter_context(Path(log).open("ab")) if log else None - process = subprocess.Popen( - [*self.base, command], - stdin=input_file, - stdout=output_file or log_file, - stderr=log_file, - env=self.environment, - ) - try: - code = process.wait(timeout=timeout) - except BaseException: - process.kill() - process.wait() - raise - if code: - raise RunPodError(f"SSH stage failed with exit code {code}; inspect remote.log") - - def upload(self, archive, workspace, output_dir, timeout): - command = f"mkdir -p {shlex.quote(workspace)} && tar xzf - --no-same-owner -C {shlex.quote(workspace)}" - self._run(command, stdin=archive, log=Path(output_dir) / "remote.log", timeout=timeout) - - def execute(self, workspace, output_dir, timeout): - self._run( - f"cd {shlex.quote(workspace)} && bash job.sh", - log=Path(output_dir) / "remote.log", - timeout=timeout, - ) - - def download(self, workspace, output_dir, timeout): - archive = Path(output_dir) / "results.tar.gz" - self._run( - f"cd {shlex.quote(workspace)}/output && tar czf - .", - stdout=archive, - log=Path(output_dir) / "remote.log", - timeout=timeout, - ) - extract_results(archive, Path(output_dir) / "results") - - -def extract_results(archive, destination): - """Accept ordinary result files only; remote tar metadata is untrusted.""" - destination = Path(destination) - destination.mkdir(parents=True, exist_ok=True) - with tarfile.open(archive, "r:gz") as handle: - members = handle.getmembers() - for member in members: - path = Path(member.name) - if path.is_absolute() or ".." in path.parts or not (member.isfile() or member.isdir()): - raise RunPodError("Downloaded results contain an unsafe tar entry") - if not (destination / path).resolve().is_relative_to(destination.resolve()): - raise RunPodError("Downloaded results escape their destination") - for member in members: - path = destination / member.name - if member.isdir(): - path.mkdir(parents=True, exist_ok=True) - else: - path.parent.mkdir(parents=True, exist_ok=True) - with path.open("wb") as output: - shutil.copyfileobj(handle.extractfile(member), output) - - -def _write_state(path, **fields): - state = json.loads(path.read_text()) if path.exists() else {} - state.update(fields) - temporary = path.with_suffix(".tmp") - temporary.write_text(json.dumps(state, indent=2) + "\n") - temporary.replace(path) - - -def _record_cleanup_state(path, **fields): - """Failure to write metadata must not prevent or misreport API cleanup.""" - try: - _write_state(path, **fields) - except (OSError, ValueError) as exc: - print(f"Could not update local state {path}: {type(exc).__name__}", flush=True) - - -def build_bundle(archive, *, mode, config_path=None, plugin_dir=None, charts=False): - if mode not in ("run", "pipeline", "check", "benchmark"): - raise ValueError("RunPod mode must be run, pipeline, check, or benchmark") - config = {} - if mode in ("run", "pipeline"): - if not config_path: - raise ValueError("RunPod run/pipeline requires --config") - config_path = Path(config_path).resolve() - config = json.loads(config_path.read_text()) - if not isinstance(config, dict) or set(config) - CONFIG_KEYS: - raise ValueError("RunPod config must contain only documented engine config keys") - if not config.get("strategy") or not config.get("input"): - raise ValueError("RunPod config must specify strategy and input") - if not isinstance(config["strategy"], str) or not all( - p.isidentifier() for p in config["strategy"].split(".") - ): - raise ValueError("RunPod strategy must be a builtin name or dotted module name") - source = (config_path.parent / config["input"]).resolve() - from .engine import validate_market_data - - validate_market_data(source, expected_interval=config.get("expected_interval")) - package = Path(__file__).parent - distribution = importlib.metadata.distribution("gpu-backtest-engine") - base_requires = [r for r in distribution.requires or [] if "extra ==" not in r] - with tempfile.TemporaryDirectory(prefix="runpod-bundle-") as directory: - staging = Path(directory) - engine = staging / "engine" - for relative in ENGINE_FILES: - path = package / relative - if path.is_symlink(): - raise ValueError("Engine bundle cannot include symlinked source files") - output = engine / "gpu_backtest" / relative - output.parent.mkdir(parents=True, exist_ok=True) - shutil.copyfile(path, output) - shutil.copyfile( - package / "runpod_requirements.txt", engine / "gpu_backtest/runpod_requirements.txt" - ) - project = ( - '[build-system]\nrequires = ["setuptools>=77", "wheel"]\nbuild-backend = "setuptools.build_meta"\n' - f'[project]\nname = "gpu-backtest-engine"\nversion = {json.dumps(__version__)}\n' - 'requires-python = ">=3.11,<3.14"\n' - f"dependencies = {json.dumps(base_requires)}\n" - '[project.scripts]\ngpu-backtest = "gpu_backtest.cli:main"\n' - '[tool.setuptools.packages.find]\nwhere = ["."]\n' - '[tool.setuptools.package-data]\ngpu_backtest = ["runpod_requirements.txt"]\n' - ) - (engine / "pyproject.toml").write_text(project) - for entry in distribution.files or []: - if str(entry).endswith("/licenses/LICENSE"): - shutil.copyfile(distribution.locate_file(entry), engine / "LICENSE") - break - if mode in ("run", "pipeline"): - (staging / "input").mkdir() - shutil.copyfile(source, staging / "input/market.csv") - config["input"] = "input/market.csv" - (staging / "job.json").write_text(json.dumps(config, indent=2) + "\n") - if plugin_dir: - plugin_dir = Path(plugin_dir).resolve() - if not plugin_dir.is_dir(): - raise ValueError("Plugin directory does not exist") - copied = 0 - for path in sorted(plugin_dir.rglob("*.py")): - relative = path.relative_to(plugin_dir) - if set(relative.parts) & IGNORED_PARTS: - continue - if path.is_symlink(): - raise ValueError("Plugin source symlinks are not supported") - output = staging / "plugins" / relative - output.parent.mkdir(parents=True, exist_ok=True) - shutil.copyfile(path, output) - copied += 1 - if not copied: - raise ValueError("Plugin directory contains no Python files") - if mode in ("run", "pipeline") and config["strategy"] not in ( - "rsi_meanrev", - "gpu_backtest.strategies.rsi_meanrev", - ): - relative = Path(*config["strategy"].split(".")) - if ( - not (staging / "plugins" / relative.with_suffix(".py")).is_file() - and not (staging / "plugins" / relative / "__init__.py").is_file() - ): - raise ValueError("External strategy must be present in --plugin-dir") - (staging / "job.sh").write_text(remote_script(mode, charts=charts)) - files = [] - for path in sorted(staging.rglob("*")): - if path.is_file(): - files.append( - { - "file": path.relative_to(staging).as_posix(), - "sha256": hashlib.sha256(path.read_bytes()).hexdigest(), - } - ) - (staging / "bundle_manifest.json").write_text(json.dumps(files, indent=2) + "\n") - with tarfile.open(archive, "w:gz") as handle: - for path in sorted(staging.rglob("*")): - if path.is_file(): - handle.add(path, arcname=path.relative_to(staging).as_posix(), recursive=False) - return files - - -def remote_script(mode, *, charts=False): - script = """#!/bin/bash -set -euo pipefail -export NUMBA_ENABLE_CUDASIM=0 -export PYTHONPATH="$PWD/plugins${PYTHONPATH:+:$PYTHONPATH}" -mkdir -p output -python3 -c 'import sys; assert (3,11) <= sys.version_info[:2] < (3,14), "Python 3.11-3.13 required"' -python3 -m venv env -env/bin/python -m pip install --disable-pip-version-check ./engine -r engine/gpu_backtest/runpod_requirements.txt -nvidia-smi --query-gpu=name,driver_version --format=csv > output/gpu.csv -env/bin/python -m gpu_backtest gpu-check --output output/gpu_check.json -""" - if mode == "run": - script += ( - "env/bin/python -m gpu_backtest run --config job.json --out-prefix output/result\n" - ) - elif mode == "pipeline": - script += "env/bin/python -m gpu_backtest pipeline --config job.json --output-dir output/pipeline\n" - elif mode == "benchmark": - script += "env/bin/python -m gpu_backtest benchmark --output output/benchmark.json --billion --billion-cpu\n" - if charts and mode in ("run", "pipeline"): - folder = "output/pipeline" if mode == "pipeline" else "output" - script += "env/bin/python -m pip install 'altair==6.3.0'\n" - script += f"env/bin/python -m gpu_backtest charts --input {folder}/*_top_entry.csv {folder}/*_top_exit.csv --output output/charts.html\n" - script += "env/bin/python -m pip freeze > output/requirements-resolved.txt\n" - return script - - -@contextlib.contextmanager -def defer_creation_signals(): - """Record Ctrl-C/termination until ownership of a creation reply is recorded.""" - interrupted = [False] - if threading.current_thread() is not threading.main_thread(): - yield interrupted - return - previous = {number: signal.getsignal(number) for number in (signal.SIGINT, signal.SIGTERM)} - - def defer(signum, frame): - interrupted[0] = True - - for number in previous: - signal.signal(number, defer) - try: - yield interrupted - finally: - for number, handler in previous.items(): - signal.signal(number, handler) - - -@contextlib.contextmanager -def terminate_on_signal(): - previous = signal.getsignal(signal.SIGTERM) - - def interrupted(signum, frame): - raise KeyboardInterrupt("RunPod launcher interrupted") - - signal.signal(signal.SIGTERM, interrupted) - try: - yield - finally: - signal.signal(signal.SIGTERM, previous) - - -def launch( - *, - config_path=None, - output_dir, - ssh_key=None, - mode="pipeline", - plugin_dir=None, - image=DEFAULT_IMAGE, - gpu="NVIDIA GeForce RTX 4090", - cloud="SECURE", - max_seconds=1800, - max_hourly_rate=1.0, - keep_pod=False, - charts=False, - dry_run=False, - client=None, - transport_factory=SSHTransport, -): - if not isinstance(max_seconds, int) or max_seconds < 1: - raise ValueError("max_seconds must be a positive integer") - if not math.isfinite(max_hourly_rate) or max_hourly_rate <= 0: - raise ValueError("max_hourly_rate must be finite and positive") - if cloud not in ("SECURE", "COMMUNITY"): - raise ValueError("cloud must be SECURE or COMMUNITY") - output_dir = Path(output_dir).resolve() - if output_dir.exists() and any(output_dir.iterdir()): - raise ValueError("RunPod output directory must be new or empty") - output_dir.mkdir(parents=True, exist_ok=True) - files = build_bundle( - output_dir / "upload.tar.gz", - mode=mode, - config_path=config_path, - plugin_dir=plugin_dir, - charts=charts, - ) - if dry_run: - print( - json.dumps( - { - "image": image, - "gpu": gpu, - "cloud": cloud, - "files": [f["file"] for f in files], - "archive": str(output_dir / "upload.tar.gz"), - "pod_created": False, - }, - indent=2, - ) - ) - return output_dir - if not ssh_key or not Path(ssh_key).expanduser().is_file(): - raise ValueError("Specify an existing --ssh-key; add its public key to RunPod Credentials") - key = Path(ssh_key).expanduser().resolve() - if not shutil.which("ssh"): - raise ValueError("An OpenSSH client is required") - client = client or RunPodClient(load_api_key()) - launch_name = "gpu-backtest-" + uuid.uuid4().hex[:12] - body = { - "name": launch_name, - "imageName": image, - "gpuTypeIds": [gpu], - "gpuCount": 1, - "cloudType": cloud, - "computeType": "GPU", - "containerDiskInGb": 30, - "volumeInGb": 0, - "ports": ["22/tcp"], - "supportPublicIp": True, - } - if mode == "benchmark": - body["minVCPUPerGPU"] = 8 - public_key = Path(str(key) + ".pub") - if public_key.is_file(): - body["env"] = {"PUBLIC_KEY": " ".join(public_key.read_text().split()[:2])} - state_path = output_dir / "runpod_state.json" - _write_state(state_path, name=launch_name, image=image, gpu=gpu, status="creating") - pod_id = None - failure = None - keep_allowed = False - print(f"Creating one {gpu} pod on {cloud}...", flush=True) - try: - with defer_creation_signals() as cancelled: - pod = client.create(body) - candidate_id = pod["id"] - if not isinstance(candidate_id, str) or not candidate_id.isalnum(): - raise RunPodError("RunPod returned an invalid pod ID") - pod_id = candidate_id - _write_state(state_path, pod_id=pod_id, status="starting") - if cancelled[0]: - raise KeyboardInterrupt("RunPod creation was cancelled") - deadline = time.monotonic() + max_seconds - quoted_rate = pod.get("adjustedCostPerHr") - if quoted_rate is None: - quoted_rate = pod.get("costPerHr") - rate = float(quoted_rate) if quoted_rate is not None else math.inf - if not math.isfinite(rate) or rate < 0 or rate > max_hourly_rate: - raise RunPodError(f"Pod rate exceeds the requested {max_hourly_rate:g}/hour limit") - keep_allowed = True - _write_state(state_path, hourly_rate=rate) - print(f"Pod {pod_id}: ${rate:g}/hour. Waiting for SSH...", flush=True) - while True: - _remaining(deadline) - details = client.get(pod_id) - connection = endpoint(details) - if connection: - transport = transport_factory(*connection, key, output_dir) - try: - if transport.ready(): - break - except subprocess.TimeoutExpired: - pass - time.sleep(min(5, _remaining(deadline))) - workspace = "/workspace/" + launch_name - _write_state(state_path, status="uploading") - print("Uploading selected engine/data/plugin files...", flush=True) - transport.upload(output_dir / "upload.tar.gz", workspace, output_dir, _remaining(deadline)) - _write_state(state_path, status="running") - print( - "Installing environment and running GPU checks/job; progress is in remote.log...", - flush=True, - ) - transport.execute(workspace, output_dir, _remaining(deadline)) - _write_state(state_path, status="downloading") - print("Downloading results...", flush=True) - transport.download(workspace, output_dir, _remaining(deadline)) - _write_state(state_path, status="completed") - except BaseException as exc: - failure = exc - _record_cleanup_state(state_path, status="failed", error=type(exc).__name__) - raise - finally: - if pod_id: - if keep_pod and keep_allowed: - _record_cleanup_state(state_path, cleanup="kept") - print( - f"Pod {pod_id} retained by --keep-pod; it continues to incur charges.", - flush=True, - ) - else: - try: - client.delete(pod_id) - except BaseException as exc: - _record_cleanup_state(state_path, cleanup="failed") - message = f"Could not delete pod {pod_id}; delete it in RunPod Console. State: {state_path}" - print(message, flush=True) - if failure is None: - raise RunPodError(message) from exc - else: - _record_cleanup_state(state_path, cleanup="deleted") - print(f"Pod {pod_id} deleted.", flush=True) - return output_dir diff --git a/src/gpu_backtest/strategies/__init__.py b/src/gpu_backtest/strategies/__init__.py deleted file mode 100644 index 30209b5..0000000 --- a/src/gpu_backtest/strategies/__init__.py +++ /dev/null @@ -1 +0,0 @@ -"""Educational strategies distributed with the engine.""" diff --git a/src/gpu_backtest/workflows/__init__.py b/src/gpu_backtest/workflows/__init__.py new file mode 100644 index 0000000..e32e536 --- /dev/null +++ b/src/gpu_backtest/workflows/__init__.py @@ -0,0 +1 @@ +"""Generic analysis workflows built on the core.""" diff --git a/src/gpu_backtest/charts.py b/src/gpu_backtest/workflows/charts.py similarity index 97% rename from src/gpu_backtest/charts.py rename to src/gpu_backtest/workflows/charts.py index 3236229..28fa509 100644 --- a/src/gpu_backtest/charts.py +++ b/src/gpu_backtest/workflows/charts.py @@ -5,7 +5,7 @@ import numpy as np import pandas as pd -from .statistics import _STAT_COLS +from gpu_backtest.core.statistics import _STAT_COLS def write_charts(input_files, output_file): diff --git a/src/gpu_backtest/common.py b/src/gpu_backtest/workflows/common.py similarity index 86% rename from src/gpu_backtest/common.py rename to src/gpu_backtest/workflows/common.py index a010dde..d3f9bf8 100644 --- a/src/gpu_backtest/common.py +++ b/src/gpu_backtest/workflows/common.py @@ -1,9 +1,7 @@ """Find parameter combinations present in every ranked top artifact.""" -import argparse import csv import os -import sys from decimal import Decimal, InvalidOperation, getcontext from pathlib import Path @@ -11,17 +9,6 @@ VALUE_COLUMN = "effect_size" -def parse_args(): - parser = argparse.ArgumentParser(description="Analyze common signals") - parser.add_argument("--input_files", nargs="+", required=True) - parser.add_argument("--key_columns", required=True, help="Comma-separated parameter columns") - parser.add_argument("--output", required=True) - parser.add_argument( - "--labels", default=None, help="Comma-separated labels matching input order" - ) - return parser.parse_args() - - def extract_label_from_filename(filename): base_name = Path(filename).stem parts = base_name.split("_") @@ -170,21 +157,3 @@ def process_data(config): print(f"Common combinations: {len(result)}") print(f"Results saved to: {output_file}") return result - - -def main(): - args = parse_args() - key_columns = [column.strip() for column in args.key_columns.split(",") if column.strip()] - labels = args.labels.split(",") if args.labels is not None else None - config = { - "input_files": args.input_files, - "key_columns": key_columns, - "output_file": args.output, - "file_labels": labels, - } - try: - process_data(config) - except (FileNotFoundError, ValueError) as exc: - print(f"ERROR: {exc}", file=sys.stderr) - return 1 - return 0 diff --git a/src/gpu_backtest/pipeline.py b/src/gpu_backtest/workflows/pipeline.py similarity index 96% rename from src/gpu_backtest/pipeline.py rename to src/gpu_backtest/workflows/pipeline.py index 7333f77..fecb458 100644 --- a/src/gpu_backtest/pipeline.py +++ b/src/gpu_backtest/workflows/pipeline.py @@ -4,10 +4,11 @@ import shutil from pathlib import Path +from gpu_backtest.core.engine import run +from gpu_backtest.core.strategy import load_strategy + from .common import process_data -from .engine import run from .splits import _read_input, split_csv_by_custom_ranges, split_csv_by_date_range -from .strategy import load_strategy def run_pipeline( diff --git a/src/gpu_backtest/splits.py b/src/gpu_backtest/workflows/splits.py similarity index 83% rename from src/gpu_backtest/splits.py rename to src/gpu_backtest/workflows/splits.py index c108a4a..ec0c20f 100644 --- a/src/gpu_backtest/splits.py +++ b/src/gpu_backtest/workflows/splits.py @@ -1,6 +1,5 @@ """Split OHLCV CSVs and record boundaries and content hashes.""" -import argparse import hashlib import json import os @@ -183,38 +182,3 @@ def split_csv_by_custom_ranges(input_file, ranges, date_column=None): specs.append((f"range_{i}", start, end_exclusive, split_df)) filtered_df = df.loc[sorted(included)] return _write_splits(input_file, date_column, len(df), filtered_df, specs, "custom_ranges") - - -def parse_args(): - parser = argparse.ArgumentParser(description="Split CSV by date range") - parser.add_argument("--input_file", required=True) - parser.add_argument("--start_date", default=None, help="Start date (DD-MM-YYYY)") - parser.add_argument("--end_date", default=None, help="End date (DD-MM-YYYY)") - parser.add_argument("--num_splits", type=int, default=3) - parser.add_argument("--date_column", default=None) - parser.add_argument( - "--ranges", - default=None, - help="Custom ranges: 'DD-MM-YYYY:DD-MM-YYYY,...'; overrides automatic splitting", - ) - return parser.parse_args() - - -def main(): - args = parse_args() - if args.ranges: - ranges = [] - for raw_range in args.ranges.split(","): - parts = raw_range.split(":") - if len(parts) != 2: - raise ValueError(f"Invalid --ranges item: {raw_range}") - ranges.append(tuple(parts)) - split_csv_by_custom_ranges(args.input_file, ranges, args.date_column) - return - df, date_column = _read_input(args.input_file, args.date_column) - start_date = args.start_date or df[date_column].min().strftime("%d-%m-%Y") - end_date = args.end_date or df[date_column].max().strftime("%d-%m-%Y") - print( - f"Configuration: {start_date} to {end_date}, {args.num_splits} splits, column {date_column}" - ) - split_csv_by_date_range(args.input_file, start_date, end_date, args.num_splits, date_column) diff --git a/tests/README.md b/tests/README.md new file mode 100644 index 0000000..03b0465 --- /dev/null +++ b/tests/README.md @@ -0,0 +1,37 @@ +# Tests: CPU, GPU simulation, and hardware + +| Directory | Runs where | What it checks | +|---|---|---| +| `cpu/` | CPU only | Core contracts, indicators/statistics, CLI/workflows, example reference, mocked RunPod lifecycle, benchmark CPU baseline, import boundaries | +| `gpu/test_simulator.py` + `simulator_cases.py` | CPU CUDA simulator in a child process | Actual GPU kernels against explicit matrix/RSI answers and CPU reference | +| `gpu/test_hardware.py` | Real NVIDIA GPU | Kernel reductions and RSI reference comparison; marked `gpu` | + +```bash +python -m pytest # CPU + isolated simulation; no rented pod +python -m pytest tests/cpu # CPU tests only +python -m pytest tests/gpu/test_simulator.py +NUMBA_ENABLE_CUDASIM=0 python -m pytest -m gpu -v +``` + +`test_simulator.py` launches `simulator_cases.py` with CUDASIM in a separate process. +The parent keeps its original environment; simulation cannot silently contaminate +later hardware checks. The case file intentionally has no `test_` prefix, so it +is collected only by the child process. Hardware tests refuse simulation and skip +if no real GPU exists; actual passes are required for hardware validation. + +Hand-calculated trades anchor timing/fees independently. A known algebraic matrix +anchors sums and float32 sums of squares. GPU kernels are compared with a Python +reference; the benchmark's parallel CPU baseline is likewise checked independently. +RunPod tests inject allocation/SSH/job/download/cleanup failure, cancellation, +price rejection, ambiguous replies, metadata failure, and unsafe archives without +leasing any resources. Architecture tests prevent examples/tools from entering core. + +The optional installed hardware smoke tool is `gpu-backtest gpu-check`; its code +is in `tools/gpu_backtest_tools/checks/`, not the engine. Performance measurements +live in [benchmark evidence](../docs/benchmarks.md). The original small hardware +validation is retained under [results](../benchmarks/results/runpod_validation_20261003.md). + +For changes to numerical logic, run both simulation and real-GPU checks. A pure +module/layout refactor must preserve the kernel body and compare previous/current +outputs. Installing the wheel and checking its RunPod export catches packaging +errors that editable-checkout tests can miss. diff --git a/tests/cpu/test_architecture.py b/tests/cpu/test_architecture.py new file mode 100644 index 0000000..4d9867e --- /dev/null +++ b/tests/cpu/test_architecture.py @@ -0,0 +1,41 @@ +"""Keep example, benchmark, and deployment dependencies out of the engine core.""" + +import ast +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[2] + + +def test_core_has_no_example_tool_cli_or_workflow_dependencies(): + for path in (ROOT / "src/gpu_backtest/core").glob("*.py"): + tree = ast.parse(path.read_text()) + for node in ast.walk(tree): + if isinstance(node, ast.ImportFrom): + module = node.module or "" + assert not module.startswith(("gpu_backtest_examples", "gpu_backtest_tools")), path + assert not module.startswith(("gpu_backtest.cli", "gpu_backtest.workflows")), path + assert node.level < 2, path + elif isinstance(node, ast.Import): + assert not any( + a.name.startswith(("gpu_backtest_examples", "gpu_backtest_tools")) + for a in node.names + ), path + + +def test_runtime_bundle_lists_only_its_explicit_packages(): + from gpu_backtest_tools.runpod.bundle import PACKAGE_FILES + + roots = { + "gpu_backtest": ROOT / "src/gpu_backtest", + "gpu_backtest_examples": ROOT / "examples/gpu_backtest_examples", + "gpu_backtest_tools": ROOT / "tools/gpu_backtest_tools", + } + assert set(PACKAGE_FILES) == set(roots) + for package, root in roots.items(): + actual = {p.relative_to(root).as_posix() for p in root.rglob("*.py")} + listed = {p for p in PACKAGE_FILES[package] if p.endswith(".py")} + assert listed == actual + assert all( + not p.startswith(("tests/", "docs/", "benchmarks/results/")) + for p in PACKAGE_FILES[package] + ) diff --git a/tests/test_benchmark.py b/tests/cpu/test_benchmark.py similarity index 81% rename from tests/test_benchmark.py rename to tests/cpu/test_benchmark.py index 96d672c..5b26428 100644 --- a/tests/test_benchmark.py +++ b/tests/cpu/test_benchmark.py @@ -4,25 +4,22 @@ import numba import numpy as np import pytest +from gpu_backtest_examples.data import generate_csv +from gpu_backtest_examples.rsi import strategy as rsi_meanrev +from gpu_backtest_tools.benchmarks import runner as benchmark +from gpu_backtest_tools.benchmarks.cpu import build_cpu_reducer +from gpu_backtest_tools.checks.reference import reference_group_sums -from gpu_backtest import benchmark -from gpu_backtest.benchmark_cpu import build_cpu_reducer -from gpu_backtest.engine import ( - decode_values, - prepare_tables, - reference_group_sums, - space_count, - validate_market_data, -) -from gpu_backtest.strategies import rsi_meanrev -from gpu_backtest.synthetic import generate_csv +from gpu_backtest.core.data import validate_market_data +from gpu_backtest.core.grid import decode_values, space_count +from gpu_backtest.core.indicators import prepare_tables def test_billion_profile_is_one_billion_unique_pairs(): entry, exit_ = benchmark.profiles()["billion"] - assert space_count(entry) == 20_000 - assert space_count(exit_) == 50_000 - assert space_count(entry) * space_count(exit_) == 1_000_000_000 + assert space_count(entry) == 20000 + assert space_count(exit_) == 50000 + assert space_count(entry) * space_count(exit_) == 1000000000 def test_parallel_cpu_baseline_matches_independent_reference(tmp_path): @@ -42,8 +39,8 @@ def test_parallel_cpu_baseline_matches_independent_reference(tmp_path): expected_entry, expected_exit = reference_group_sums( rsi_meanrev, *arrays, entry, exit_, 0.0015, 0.0015 ) - np.testing.assert_allclose(results[0], expected_entry, rtol=0, atol=1e-5) - np.testing.assert_allclose(results[2], expected_exit, rtol=0, atol=1e-5) + np.testing.assert_allclose(results[0], expected_entry, rtol=0, atol=1e-05) + np.testing.assert_allclose(results[2], expected_exit, rtol=0, atol=1e-05) specs = SimpleNamespace(ENTRY_DIMS=entry, EXIT_DIMS=exit_, TABLES=rsi_meanrev.TABLES) tables, maps = prepare_tables(specs, *arrays) matrix = np.empty((space_count(entry), space_count(exit_)), np.float32) @@ -69,7 +66,9 @@ def test_benchmark_rejects_simulator_and_missing_hardware(monkeypatch, tmp_path) def test_generated_data_reproduces_public_example(tmp_path): generated = generate_csv(tmp_path / "bars.csv", 128) - original = Path(__file__).resolve().parents[1] / "examples/synthetic.csv" + original = ( + Path(__file__).resolve().parents[2] / "examples/gpu_backtest_examples/rsi/synthetic.csv" + ) assert generated.read_bytes() == original.read_bytes() @@ -78,7 +77,7 @@ def test_cpu_baseline_writes_complete_ranked_artifacts(tmp_path): source = generate_csv(tmp_path / "bars.csv", 32) prefix = tmp_path / "cpu" - entry, exit_ = rsi_meanrev.ENTRY_DIMS, rsi_meanrev.EXIT_DIMS + entry, exit_ = (rsi_meanrev.ENTRY_DIMS, rsi_meanrev.EXIT_DIMS) old = numba.get_num_threads() numba.set_num_threads(min(2, numba.config.NUMBA_NUM_THREADS)) try: diff --git a/tests/test_cli.py b/tests/cpu/test_cli.py similarity index 84% rename from tests/test_cli.py rename to tests/cpu/test_cli.py index 8d91dbb..cf365e6 100644 --- a/tests/test_cli.py +++ b/tests/cpu/test_cli.py @@ -7,7 +7,7 @@ import pandas as pd import pytest -ROOT = Path(__file__).resolve().parents[1] +ROOT = Path(__file__).resolve().parents[2] def invoke(*args, simulation=False, cwd=None): @@ -29,7 +29,7 @@ def test_config_paths_resolve_outside_checkout(tmp_path): result = invoke( "run", "--config", - ROOT / "examples/rsi.json", + ROOT / "examples/gpu_backtest_examples/rsi/config.json", "--out-prefix", prefix, simulation=True, @@ -42,7 +42,11 @@ def test_config_paths_resolve_outside_checkout(tmp_path): def test_strategy_is_always_explicit(tmp_path): result = invoke( - "run", "--input", ROOT / "examples/synthetic.csv", "--out-prefix", tmp_path / "x" + "run", + "--input", + ROOT / "examples/gpu_backtest_examples/rsi/synthetic.csv", + "--out-prefix", + tmp_path / "x", ) assert result.returncode == 2 and "explicitly" in result.stderr assert not (tmp_path / "x_top_entry.csv").exists() @@ -51,7 +55,12 @@ def test_strategy_is_always_explicit(tmp_path): def test_complete_public_pipeline(tmp_path): output = tmp_path / "pipeline" result = invoke( - "pipeline", "--config", ROOT / "examples/rsi.json", "--output-dir", output, simulation=True + "pipeline", + "--config", + ROOT / "examples/gpu_backtest_examples/rsi/config.json", + "--output-dir", + output, + simulation=True, ) assert result.returncode == 0, result.stdout + result.stderr manifest = json.loads((output / "split_manifest.json").read_text()) diff --git a/tests/test_common_contract.py b/tests/cpu/test_common_contract.py similarity index 100% rename from tests/test_common_contract.py rename to tests/cpu/test_common_contract.py diff --git a/tests/test_engine_contract.py b/tests/cpu/test_engine_contract.py similarity index 81% rename from tests/test_engine_contract.py rename to tests/cpu/test_engine_contract.py index 2c0fe0a..78275f3 100644 --- a/tests/test_engine_contract.py +++ b/tests/cpu/test_engine_contract.py @@ -4,7 +4,8 @@ import pandas as pd import pytest -from gpu_backtest import engine as harness +from gpu_backtest.core import data as market_data +from gpu_backtest.core import output as artifacts def _write_market_csv(path, periods=8): @@ -24,7 +25,7 @@ def _write_market_csv(path, periods=8): def test_market_data_validation_accepts_explicit_interval(tmp_path): source = tmp_path / "market.csv" _write_market_csv(source) - data = harness.validate_market_data(source, expected_interval="12h") + data = market_data.validate_market_data(source, expected_interval="12h") assert len(data) == 8 @@ -34,18 +35,18 @@ def test_market_data_validation_rejects_gap_and_bad_ohlc(tmp_path): data = pd.read_csv(source).drop(index=4).reset_index(drop=True) data.to_csv(source, index=False) with pytest.raises(ValueError, match="bar interval"): - harness.validate_market_data(source, expected_interval="12h") + market_data.validate_market_data(source, expected_interval="12h") _write_market_csv(source) data = pd.read_csv(source) data.loc[2, "High"] = data.loc[2, "Low"] - 1 data.to_csv(source, index=False) with pytest.raises(ValueError, match="High is below"): - harness.validate_market_data(source, expected_interval="12h") + market_data.validate_market_data(source, expected_interval="12h") def test_reduction_contract_checks_sum_and_sumsq(): matrix = np.array([[1.0, 2.0], [3.0, 4.0]], dtype=np.float32) - checks = harness.validate_reduction_outputs( + checks = artifacts.validate_reduction_outputs( matrix.sum(axis=1, dtype=np.float64), (matrix * matrix).sum(axis=1, dtype=np.float64), matrix.sum(axis=0, dtype=np.float64), @@ -54,7 +55,7 @@ def test_reduction_contract_checks_sum_and_sumsq(): assert checks["sum"]["difference"] == 0 assert checks["sumsq"]["difference"] == 0 with pytest.raises(RuntimeError, match="grand sum mismatch"): - harness.validate_reduction_outputs([3.0, 7.0], [5.0, 25.0], [4.0, 7.0], [10.0, 20.0]) + artifacts.validate_reduction_outputs([3.0, 7.0], [5.0, 25.0], [4.0, 7.0], [10.0, 20.0]) def test_top_artifact_preserves_float64_effects_and_writes_manifest(tmp_path): @@ -63,11 +64,11 @@ def test_top_artifact_preserves_float64_effects_and_writes_manifest(tmp_path): entry_sumsq = (matrix * matrix).sum(axis=1, dtype=np.float64) exit_sum = matrix.sum(axis=0, dtype=np.float64) exit_sumsq = (matrix * matrix).sum(axis=0, dtype=np.float64) - checks = harness.validate_reduction_outputs(entry_sum, entry_sumsq, exit_sum, exit_sumsq) + checks = artifacts.validate_reduction_outputs(entry_sum, entry_sumsq, exit_sum, exit_sumsq) prefix = str(tmp_path / "result") entry_dims = [("entry_parameter", 0.1, 0.2, 0.1, True)] exit_dims = [("exit_parameter", 1, 2, 1, False)] - harness._write_top_csv( + artifacts.write_top_csv( prefix, "test_strategy", "source.csv", @@ -81,7 +82,7 @@ def test_top_artifact_preserves_float64_effects_and_writes_manifest(tmp_path): checks, False, ) - expected = harness.cs.stats_from_groups( + expected = artifacts.cs.stats_from_groups( entry_sum, entry_sumsq, 2, entry_sum.sum(), entry_sumsq.sum(), 4 )["effect_size"] output = pd.read_csv(f"{prefix}_top_entry.csv") diff --git a/tests/cpu/test_examples.py b/tests/cpu/test_examples.py new file mode 100644 index 0000000..a40f737 --- /dev/null +++ b/tests/cpu/test_examples.py @@ -0,0 +1,34 @@ +"""Example-specific CPU truth lives outside generic engine tests.""" + +import numpy as np +import pytest +from gpu_backtest_examples.rsi import strategy + +from gpu_backtest.core.indicators import prepare_tables + + +@pytest.mark.parametrize("fee,expected", [(0.0, 50.0), (0.0015, 49.625)]) +def test_rsi_cpu_hand_calculated_trade(fee, expected): + bars = [ + (100, 101, 99, 100, 10), + (95, 96, 89, 90, 10), + (80, 81, 79, 80, 10), + (90, 101, 90, 100, 10), + (110, 121, 110, 120, 10), + (120, 121, 119, 120, 10), + (120, 121, 119, 120, 10), + ] + raw = np.array(bars, dtype=np.float32) + arrays = [raw[:, i] for i in (3, 0, 1, 2, 4)] + from types import SimpleNamespace + + dims = SimpleNamespace( + ENTRY_DIMS=[("p_e", 2, 2, 1, False)], + EXIT_DIMS=[("p_x", 2, 2, 1, False)], + TABLES=strategy.TABLES, + ) + tables, maps = prepare_tables(dims, *arrays) + result = strategy.reference( + *arrays, tables, maps, [2.0, 30.0, 0.0, 0.0], [2.0, 70.0, 0.0, 0.0], fee, fee + ) + assert result == pytest.approx(expected, abs=1e-6) diff --git a/tests/test_gpu_check.py b/tests/cpu/test_gpu_check.py similarity index 91% rename from tests/test_gpu_check.py rename to tests/cpu/test_gpu_check.py index 92a311c..6b2dc63 100644 --- a/tests/test_gpu_check.py +++ b/tests/cpu/test_gpu_check.py @@ -1,8 +1,7 @@ from types import SimpleNamespace import pytest - -from gpu_backtest import gpu_check +from gpu_backtest_tools.checks import gpu as gpu_check def test_gpu_check_refuses_simulation(monkeypatch): diff --git a/tests/test_runpod.py b/tests/cpu/test_runpod.py similarity index 90% rename from tests/test_runpod.py rename to tests/cpu/test_runpod.py index 895512d..4f17c95 100644 --- a/tests/test_runpod.py +++ b/tests/cpu/test_runpod.py @@ -7,10 +7,11 @@ from pathlib import Path import pytest +from gpu_backtest_tools.runpod import client as runpod_client +from gpu_backtest_tools.runpod import launcher as runpod +from gpu_backtest_tools.runpod import transport as runpod_transport -from gpu_backtest import runpod - -ROOT = Path(__file__).resolve().parents[1] +ROOT = Path(__file__).resolve().parents[2] class Client: @@ -65,7 +66,7 @@ def launch_options(tmp_path, client): key = tmp_path / "ssh_key" key.write_text("test-key") return dict( - config_path=ROOT / "examples/rsi.json", + config_path=ROOT / "examples/gpu_backtest_examples/rsi/config.json", output_dir=tmp_path / "output", ssh_key=key, client=client, @@ -156,6 +157,7 @@ def test_rate_guard_deletes_even_with_keep_pod(tmp_path, rate): def test_cancel_during_creation_records_ownership_before_cleanup(tmp_path): + class CancelClient(Client): def create(self, body): os.kill(os.getpid(), signal.SIGINT) @@ -169,13 +171,14 @@ def create(self, body): def test_dry_run_never_reads_credentials_or_creates_a_pod(tmp_path, monkeypatch): + def forbidden(): raise AssertionError("Credentials must not be read") monkeypatch.setattr(runpod, "load_api_key", forbidden) client = Client() runpod.launch( - config_path=ROOT / "examples/rsi.json", + config_path=ROOT / "examples/gpu_backtest_examples/rsi/config.json", output_dir=tmp_path / "output", dry_run=True, client=client, @@ -187,7 +190,7 @@ def forbidden(): def test_benchmark_bundle_needs_no_config_and_runs_billion_profile(tmp_path): archive = tmp_path / "benchmark.tar.gz" files = runpod.build_bundle(archive, mode="benchmark") - assert "engine/gpu_backtest/benchmark.py" in {f["file"] for f in files} + assert "engine/gpu_backtest_tools/benchmarks/runner.py" in {f["file"] for f in files} with tarfile.open(archive) as bundle: command = bundle.extractfile("job.sh").read().decode() assert "benchmark --output output/benchmark.json --billion --billion-cpu" in command @@ -203,8 +206,11 @@ def test_bundle_contains_only_selected_python_plugin_and_data(tmp_path): (plugins / ".git/unrelated.py").write_text("do not upload") (plugins / ".env").write_text("do not upload") (plugins / "ssh_key").write_text("do not upload") - config = json.loads((ROOT / "examples/rsi.json").read_text()) - config.update(strategy="custom.strategy", input=str(ROOT / "examples/synthetic.csv")) + config = json.loads((ROOT / "examples/gpu_backtest_examples/rsi/config.json").read_text()) + config.update( + strategy="custom.strategy", + input=str(ROOT / "examples/gpu_backtest_examples/rsi/synthetic.csv"), + ) path = tmp_path / "job.json" path.write_text(json.dumps(config)) archive = tmp_path / "bundle.tar.gz" @@ -213,7 +219,7 @@ def test_bundle_contains_only_selected_python_plugin_and_data(tmp_path): assert {"plugins/custom/strategy.py", "plugins/custom/__init__.py", "engine/LICENSE"}.issubset( names ) - assert not any(".git" in n or ".env" in n or "ssh_key" in n for n in names) + assert not any((".git" in n or ".env" in n or "ssh_key" in n for n in names)) with tarfile.open(archive) as handle: normalized = json.load(handle.extractfile("job.json")) assert normalized["input"] == "input/market.csv" @@ -222,8 +228,11 @@ def test_bundle_contains_only_selected_python_plugin_and_data(tmp_path): def test_missing_plugin_and_unknown_config_fields_fail_before_creation(tmp_path): path = tmp_path / "job.json" - config = json.loads((ROOT / "examples/rsi.json").read_text()) - config.update(strategy="missing.strategy", input=str(ROOT / "examples/synthetic.csv")) + config = json.loads((ROOT / "examples/gpu_backtest_examples/rsi/config.json").read_text()) + config.update( + strategy="missing.strategy", + input=str(ROOT / "examples/gpu_backtest_examples/rsi/synthetic.csv"), + ) path.write_text(json.dumps(config)) with pytest.raises(ValueError, match="--plugin-dir"): runpod.build_bundle(tmp_path / "upload.tar.gz", mode="run", config_path=path) @@ -257,7 +266,7 @@ def test_download_rejects_path_traversal_and_links(tmp_path, name, kind): entry.size = 1 handle.addfile(entry, io.BytesIO(b"x")) with pytest.raises(runpod.RunPodError, match="unsafe tar"): - runpod.extract_results(archive, tmp_path / "results") + runpod_transport.extract_results(archive, tmp_path / "results") assert not (tmp_path / "escape").exists() @@ -271,7 +280,7 @@ def fail(request, timeout): request.full_url, 403, "Forbidden", {}, io.BytesIO(key.encode()) ) - monkeypatch.setattr(runpod.urllib.request, "urlopen", fail) + monkeypatch.setattr(runpod_client.urllib.request, "urlopen", fail) with pytest.raises(runpod.RunPodError) as error: runpod.RunPodClient(key).request("GET", "/pods") assert key not in str(error.value) diff --git a/tests/test_split_contract.py b/tests/cpu/test_split_contract.py similarity index 97% rename from tests/test_split_contract.py rename to tests/cpu/test_split_contract.py index 80ac1a9..2372e01 100644 --- a/tests/test_split_contract.py +++ b/tests/cpu/test_split_contract.py @@ -3,7 +3,7 @@ import pandas as pd import pytest -from gpu_backtest import splits as splitter +from gpu_backtest.workflows import splits as splitter def _market_frame(periods=10): diff --git a/tests/test_statistics.py b/tests/cpu/test_statistics.py similarity index 98% rename from tests/test_statistics.py rename to tests/cpu/test_statistics.py index ac9d46b..1cfcf3b 100644 --- a/tests/test_statistics.py +++ b/tests/cpu/test_statistics.py @@ -2,7 +2,7 @@ import numpy as np -from gpu_backtest import statistics as cs +from gpu_backtest.core import statistics as cs def _effect_size_scalar(sum1, sumsq1, n1, tot_sum, tot_sumsq, tot_count): diff --git a/tests/test_strategy.py b/tests/cpu/test_strategy.py similarity index 84% rename from tests/test_strategy.py rename to tests/cpu/test_strategy.py index f54f417..18e48ce 100644 --- a/tests/test_strategy.py +++ b/tests/cpu/test_strategy.py @@ -2,8 +2,9 @@ import pytest -from gpu_backtest.engine import decode_values, run -from gpu_backtest.strategy import load_strategy, validate_strategy +from gpu_backtest.core.engine import run +from gpu_backtest.core.grid import decode_values +from gpu_backtest.core.strategy import load_strategy, validate_strategy def specification(**changes): @@ -19,9 +20,9 @@ def specification(**changes): def test_builtin_and_module_object_loading(): - strategy = load_strategy("rsi_meanrev") + strategy = load_strategy("gpu_backtest_examples.rsi.strategy") assert load_strategy(strategy) is strategy - assert load_strategy("gpu_backtest.strategies.rsi_meanrev") is strategy + assert load_strategy("gpu_backtest_examples.rsi.strategy") is strategy validate_strategy(strategy) @@ -74,9 +75,9 @@ def test_windows_must_be_positive_integer_dimensions(): def test_parameter_decode_preserves_sub_micro_precision(): - dims = [("fraction", 0.0000001, 0.0000002, 0.0000001, True)] - assert decode_values(dims, 0) == [0.0000001] - assert decode_values(dims, 1) == [0.0000002] + dims = [("fraction", 1e-07, 2e-07, 1e-07, True)] + assert decode_values(dims, 0) == [1e-07] + assert decode_values(dims, 1) == [2e-07] @pytest.mark.parametrize("value", ["", "../strategy", "thing:callback"]) diff --git a/tests/test_tables.py b/tests/cpu/test_tables.py similarity index 97% rename from tests/test_tables.py rename to tests/cpu/test_tables.py index dd7505f..9497445 100644 --- a/tests/test_tables.py +++ b/tests/cpu/test_tables.py @@ -1,6 +1,6 @@ import numpy as np -from gpu_backtest.tables import build_table +from gpu_backtest.core.indicators import build_table def _bt(kind, source, windows, close, open_=None, high=None, low=None, volume=None): diff --git a/tests/helpers.py b/tests/gpu/fixtures.py similarity index 81% rename from tests/helpers.py rename to tests/gpu/fixtures.py index 1b152ac..6c5988e 100644 --- a/tests/helpers.py +++ b/tests/gpu/fixtures.py @@ -4,9 +4,10 @@ import numpy as np import pandas as pd +from gpu_backtest_tools.checks.reference import reference_group_sums -from gpu_backtest.engine import reference_group_sums, run -from gpu_backtest.strategy import load_strategy +from gpu_backtest import run +from gpu_backtest.core.strategy import load_strategy def write_bars(path, bars): @@ -32,24 +33,10 @@ def matrix_plugin(tmp_path, monkeypatch): package = tmp_path / "outside_plugins" package.mkdir() (package / "__init__.py").write_text("") - (package / "matrix.py").write_text(""" -NAME = "matrix" -ENTRY_DIMS = [("entry", 1, 2, 1, False)] -EXIT_DIMS = [("exit", 1, 3, 1, False)] -TABLES = [] - -def make_device_fn(cuda): - @cuda.jit(device=True, inline=True) - def algo(close, open_, high, low, vol, t0, t1, t2, t3, - rm0, rm1, rm2, rm3, ep, xp, num_bars, buy, sell): - return ep[0] * 10 + xp[0] + 1 - return algo - -def reference(close, open_, high, low, vol, tabs, rms, ep, xp, buy, sell): - return ep[0] * 10 + xp[0] + 1 -""") + (package / "matrix.py").write_text( + '\nNAME = "matrix"\nENTRY_DIMS = [("entry", 1, 2, 1, False)]\nEXIT_DIMS = [("exit", 1, 3, 1, False)]\nTABLES = []\n\ndef make_device_fn(cuda):\n @cuda.jit(device=True, inline=True)\n def algo(close, open_, high, low, vol, t0, t1, t2, t3,\n rm0, rm1, rm2, rm3, ep, xp, num_bars, buy, sell):\n return ep[0] * 10 + xp[0] + 1\n return algo\n\ndef reference(close, open_, high, low, vol, tabs, rms, ep, xp, buy, sell):\n return ep[0] * 10 + xp[0] + 1\n' + ) monkeypatch.syspath_prepend(str(tmp_path)) - # A fresh package directory is used for each case, including repeated runs. import sys sys.modules.pop("outside_plugins.matrix", None) @@ -91,7 +78,7 @@ def check_matrix(tmp_path, monkeypatch): def check_rsi_reference(tmp_path): source = tmp_path / "bars.csv" data = write_bars(source, rsi_bars()) - strategy = load_strategy("rsi_meanrev") + strategy = load_strategy("gpu_backtest_examples.rsi.strategy") result = run(strategy, source, None, verbose=False) args = [ data[column].to_numpy(dtype=np.float32) @@ -100,15 +87,15 @@ def check_rsi_reference(tmp_path): entry, exit_ = reference_group_sums( strategy, *args, strategy.ENTRY_DIMS, strategy.EXIT_DIMS, 0.0015, 0.0015 ) - np.testing.assert_allclose(result["entry_sum"], entry, rtol=1e-5, atol=1e-4) - np.testing.assert_allclose(result["exit_sum"], exit_, rtol=1e-5, atol=1e-4) + np.testing.assert_allclose(result["entry_sum"], entry, rtol=1e-05, atol=0.0001) + np.testing.assert_allclose(result["exit_sum"], exit_, rtol=1e-05, atol=0.0001) assert np.abs(result["entry_sum"]).max() > 1 def single_rsi(path, bars, *, buy=0.0, sell=0.0, buy_level=30): write_bars(path, bars) result = run( - "rsi_meanrev", + "gpu_backtest_examples.rsi.strategy", path, None, buy=buy, diff --git a/tests/simulator_cases.py b/tests/gpu/simulator_cases.py similarity index 79% rename from tests/simulator_cases.py rename to tests/gpu/simulator_cases.py index 202fe54..afc9f7c 100644 --- a/tests/simulator_cases.py +++ b/tests/gpu/simulator_cases.py @@ -1,7 +1,7 @@ """Explicitly collected only in an isolated CUDA-simulator process.""" import pytest -from helpers import check_matrix, check_rsi_reference, rsi_bars, single_rsi +from fixtures import check_matrix, check_rsi_reference, rsi_bars, single_rsi def test_known_matrix_external_module_and_determinism(tmp_path, monkeypatch): @@ -13,7 +13,7 @@ def test_rsi_matches_cpu_reference(tmp_path): def test_packaged_gpu_check_numeric_anchors(tmp_path): - from gpu_backtest.gpu_check import check_numerics + from gpu_backtest_tools.checks.gpu import check_numerics assert check_numerics(tmp_path) == { "known_matrix": "pass", @@ -24,16 +24,11 @@ def test_packaged_gpu_check_numeric_anchors(tmp_path): @pytest.mark.parametrize( "exit_open,fee,expected", - [ - (120, 0.0, 50.0), - (120, 0.0015, 49.625), - (60, 0.0, -25.0), - (60, 0.0015, -25.2625), - ], + [(120, 0.0, 50.0), (120, 0.0015, 49.625), (60, 0.0, -25.0), (60, 0.0015, -25.2625)], ) def test_rsi_hand_calculated_round_trip(tmp_path, exit_open, fee, expected): actual = single_rsi(tmp_path / "bars.csv", rsi_bars(exit_open), buy=fee, sell=fee) - assert actual == pytest.approx(expected, abs=1e-4) + assert actual == pytest.approx(expected, abs=0.0001) def test_no_entry_returns_zero(tmp_path): @@ -59,4 +54,4 @@ def test_open_position_is_marked_at_final_open_without_sell_fee(tmp_path): (60, 71, 59, 70, 10), ] actual = single_rsi(tmp_path / "bars.csv", bars, buy=0.0015, sell=0.5) - assert actual == pytest.approx(-25.15, abs=1e-4) + assert actual == pytest.approx(-25.15, abs=0.0001) diff --git a/tests/test_gpu.py b/tests/gpu/test_hardware.py similarity index 89% rename from tests/test_gpu.py rename to tests/gpu/test_hardware.py index 34b5c0e..2b53095 100644 --- a/tests/test_gpu.py +++ b/tests/gpu/test_hardware.py @@ -1,5 +1,5 @@ import pytest -from helpers import check_matrix, check_rsi_reference +from fixtures import check_matrix, check_rsi_reference from numba import config, cuda pytestmark = pytest.mark.gpu diff --git a/tests/test_simulator.py b/tests/gpu/test_simulator.py similarity index 100% rename from tests/test_simulator.py rename to tests/gpu/test_simulator.py diff --git a/tools/gpu_backtest_tools/__init__.py b/tools/gpu_backtest_tools/__init__.py new file mode 100644 index 0000000..7839048 --- /dev/null +++ b/tools/gpu_backtest_tools/__init__.py @@ -0,0 +1 @@ +"""Optional execution, benchmark, and hardware-check tools.""" diff --git a/tools/gpu_backtest_tools/benchmarks/__init__.py b/tools/gpu_backtest_tools/benchmarks/__init__.py new file mode 100644 index 0000000..80f199a --- /dev/null +++ b/tools/gpu_backtest_tools/benchmarks/__init__.py @@ -0,0 +1 @@ +"""RSI performance measurement and its CPU comparison baseline.""" diff --git a/src/gpu_backtest/benchmark_cpu.py b/tools/gpu_backtest_tools/benchmarks/cpu.py similarity index 97% rename from src/gpu_backtest/benchmark_cpu.py rename to tools/gpu_backtest_tools/benchmarks/cpu.py index bd1dbbe..4148031 100644 --- a/src/gpu_backtest/benchmark_cpu.py +++ b/tools/gpu_backtest_tools/benchmarks/cpu.py @@ -74,6 +74,6 @@ def reduce( squared += np.float64(np.float32(result * result)) exit_sum[x] = total exit_sumsq[x] = squared - return entry_sum, entry_sumsq, exit_sum, exit_sumsq + return (entry_sum, entry_sumsq, exit_sum, exit_sumsq) return reduce diff --git a/src/gpu_backtest/benchmark.py b/tools/gpu_backtest_tools/benchmarks/runner.py similarity index 89% rename from src/gpu_backtest/benchmark.py rename to tools/gpu_backtest_tools/benchmarks/runner.py index 0567dfa..50a44f6 100644 --- a/src/gpu_backtest/benchmark.py +++ b/tools/gpu_backtest_tools/benchmarks/runner.py @@ -14,22 +14,20 @@ import numba import numpy as np +from gpu_backtest_examples.data import generate_csv +from gpu_backtest_examples.rsi import strategy as rsi_meanrev from numba import config, cuda -from .benchmark_cpu import build_cpu_reducer -from .engine import ( - _write_top_csv, - build_kernels, - dim_arrays, - prepare_tables, - run, - space_count, - validate_market_data, - validate_reduction_outputs, -) -from .strategies import rsi_meanrev -from .strategy import validate_strategy -from .synthetic import generate_csv +from gpu_backtest.core.data import validate_market_data +from gpu_backtest.core.engine import run +from gpu_backtest.core.grid import dim_arrays, space_count +from gpu_backtest.core.indicators import prepare_tables +from gpu_backtest.core.kernels import build_kernels +from gpu_backtest.core.output import validate_reduction_outputs +from gpu_backtest.core.output import write_top_csv as _write_top_csv +from gpu_backtest.core.strategy import validate_strategy + +from .cpu import build_cpu_reducer RESULT_KEYS = ("entry_sum", "entry_sumsq", "exit_sum", "exit_sumsq") @@ -77,9 +75,9 @@ class GPUPlan: """Prepare and compile actual engine kernels once, then time both reductions.""" def __init__(self, strategy, args, threads_per_block=128): - arrays, tabs, maps = args[:5], args[5], args[6] + arrays, tabs, maps = (args[:5], args[5], args[6]) e_lo, e_step, e_count, n_e, x_lo, x_step, x_count, n_x, entries, exits, buy, sell = args[7:] - self.entries, self.exits = entries, exits + self.entries, self.exits = (entries, exits) self.tpb = threads_per_block self.kernels = build_kernels(strategy) self.outputs = [ @@ -111,7 +109,7 @@ def execute(self): cuda.synchronize() def results(self): - return tuple(array.copy_to_host() for array in self.outputs) + return tuple((array.copy_to_host() for array in self.outputs)) def _timed(call, repeats): @@ -121,7 +119,7 @@ def _timed(call, repeats): started = time.perf_counter() result = call() values.append(time.perf_counter() - started) - return result, values + return (result, values) def _cpu_info(threads): @@ -176,11 +174,11 @@ def benchmark(output, *, bars=1024, cpu_threads=8, repeats=3, billion=False, bil raise ValueError("Performance benchmarking requires real CUDA, not CUDASIM") if not cuda.is_available(): raise ValueError("Performance benchmarking requires an available NVIDIA GPU") - if not isinstance(repeats, int) or repeats < 1 or not isinstance(bars, int) or bars < 2: + if not isinstance(repeats, int) or repeats < 1 or (not isinstance(bars, int)) or (bars < 2): raise ValueError("Require positive repeats and at least two bars") if not isinstance(cpu_threads, int) or not 1 <= cpu_threads <= numba.config.NUMBA_NUM_THREADS: raise ValueError("CPU thread count exceeds the available Numba thread pool") - if billion_cpu and not billion: + if billion_cpu and (not billion): raise ValueError("--billion-cpu requires --billion") output = Path(output) output.parent.mkdir(parents=True, exist_ok=True) @@ -213,12 +211,12 @@ def benchmark(output, *, bars=1024, cpu_threads=8, repeats=3, billion=False, bil gpu_results = gpu.results() differences = {} for key, left, right in zip(RESULT_KEYS, cpu_results, gpu_results): - np.testing.assert_allclose(left, right, rtol=1e-6, atol=1e-3) + np.testing.assert_allclose(left, right, rtol=1e-06, atol=0.001) differences[key] = float(np.max(np.abs(left - right))) validate_reduction_outputs(*gpu_results) - if not np.any(np.abs(gpu_results[0]) > 1e-6): + if not np.any(np.abs(gpu_results[0]) > 1e-06): raise RuntimeError("Benchmark strategy produced only zero returns") - cpu_median, gpu_median = statistics.median(cpu_times), statistics.median(gpu_times) + cpu_median, gpu_median = (statistics.median(cpu_times), statistics.median(gpu_times)) device = cuda.get_current_device() name = device.name.decode() if isinstance(device.name, bytes) else str(device.name) packages = {} @@ -278,13 +276,13 @@ def benchmark(output, *, bars=1024, cpu_threads=8, repeats=3, billion=False, bil ) if billion: e_dims, x_dims = profiles()["billion"] - entries, exits = space_count(e_dims), space_count(x_dims) - if entries * exits != 1_000_000_000: + entries, exits = (space_count(e_dims), space_count(x_dims)) + if entries * exits != 1000000000: raise RuntimeError("Billion profile must contain exactly one billion pairs") print("Running the normal engine API on 1,000,000,000 unique pairs...", flush=True) started = time.perf_counter() large = run( - "rsi_meanrev", + "gpu_backtest_examples.rsi.strategy", source, str(output.parent / "billion"), entry_dims=e_dims, @@ -301,7 +299,7 @@ def benchmark(output, *, bars=1024, cpu_threads=8, repeats=3, billion=False, bil "unique_combinations": entries * exits, "strategy_evaluations": 2 * entries * exits, "engine_end_to_end_seconds": elapsed, - "reduction_array_bytes": sum(a.nbytes for a in large.values()), + "reduction_array_bytes": sum((a.nbytes for a in large.values())), "hypothetical_float32_matrix_bytes": entries * exits * 4, "nonzero_entry_groups": int(np.count_nonzero(large["entry_sum"])), "timing_scope": "Normal engine run including CSV validation, indicator tables, allocation/transfers, kernel construction/JIT, two GPU passes, copies, statistics and top CSV/manifest output. Excludes pod provisioning and dependency installation.", @@ -320,7 +318,7 @@ def benchmark(output, *, bars=1024, cpu_threads=8, repeats=3, billion=False, bil cpu_elapsed = time.perf_counter() - started differences = {} for key in RESULT_KEYS: - np.testing.assert_allclose(cpu_large[key], large[key], rtol=1e-6, atol=1e-3) + np.testing.assert_allclose(cpu_large[key], large[key], rtol=1e-06, atol=0.001) differences[key] = float(np.max(np.abs(cpu_large[key] - large[key]))) report["billion"].update( cpu_engine_end_to_end_seconds=cpu_elapsed, diff --git a/tools/gpu_backtest_tools/checks/__init__.py b/tools/gpu_backtest_tools/checks/__init__.py new file mode 100644 index 0000000..cb2a70e --- /dev/null +++ b/tools/gpu_backtest_tools/checks/__init__.py @@ -0,0 +1 @@ +"""Numerical references and real-GPU smoke checks.""" diff --git a/src/gpu_backtest/gpu_check.py b/tools/gpu_backtest_tools/checks/gpu.py similarity index 92% rename from src/gpu_backtest/gpu_check.py rename to tools/gpu_backtest_tools/checks/gpu.py index bf906e7..628fe2a 100644 --- a/src/gpu_backtest/gpu_check.py +++ b/tools/gpu_backtest_tools/checks/gpu.py @@ -10,10 +10,11 @@ import pandas as pd from numba import config, cuda -from .engine import run +from gpu_backtest.core.engine import run def _matrix_device(cuda): + @cuda.jit(device=True, inline=True) def algo( close, @@ -80,10 +81,17 @@ def check_numerics(directory): x = [("p_x", 2, 2, 1, False), ("sell_lvl", 70, 70, 1, False)] for fee, expected_return in [(0.0, 50.0), (0.0015, 49.625)]: result = run( - "rsi_meanrev", source, None, buy=fee, sell=fee, entry_dims=e, exit_dims=x, verbose=False + "gpu_backtest_examples.rsi.strategy", + source, + None, + buy=fee, + sell=fee, + entry_dims=e, + exit_dims=x, + verbose=False, ) - np.testing.assert_allclose(result["entry_sum"], [expected_return], rtol=0, atol=1e-4) - np.testing.assert_allclose(result["exit_sum"], [expected_return], rtol=0, atol=1e-4) + np.testing.assert_allclose(result["entry_sum"], [expected_return], rtol=0, atol=0.0001) + np.testing.assert_allclose(result["exit_sum"], [expected_return], rtol=0, atol=0.0001) return {"known_matrix": "pass", "repeat_determinism": "pass", "rsi_hand_calculation": "pass"} diff --git a/tools/gpu_backtest_tools/checks/reference.py b/tools/gpu_backtest_tools/checks/reference.py new file mode 100644 index 0000000..30c25c9 --- /dev/null +++ b/tools/gpu_backtest_tools/checks/reference.py @@ -0,0 +1,32 @@ +"""Small-grid CPU oracle used by tests and hardware checks, never by the GPU runner.""" + +import numpy as np + +from gpu_backtest.core.grid import decode_values, space_count +from gpu_backtest.core.indicators import prepare_tables +from gpu_backtest.core.strategy import MAX_DIMS, validate_strategy + + +def reference_group_sums(strategy, close, open_, high, low, vol, e_dims, x_dims, buy, sell): + """Compute small-grid grouped returns with a strategy CPU reference.""" + strategy = validate_strategy(strategy, e_dims, x_dims) + if not callable(getattr(strategy, "reference", None)): + raise ValueError("CPU comparisons require a callable strategy.reference") + + class _S: + ENTRY_DIMS, EXIT_DIMS, TABLES = (e_dims, x_dims, strategy.TABLES) + + tabs, rms = prepare_tables(_S, close, open_, high, low, vol) + ec, xc = (space_count(e_dims), space_count(x_dims)) + es = np.zeros(ec) + xs = np.zeros(xc) + for e in range(ec): + ep = decode_values(e_dims, e) + [0.0] * (MAX_DIMS - len(e_dims)) + for x in range(xc): + xp = decode_values(x_dims, x) + [0.0] * (MAX_DIMS - len(x_dims)) + tr = np.float32( + strategy.reference(close, open_, high, low, vol, tabs, rms, ep, xp, buy, sell) + ) + es[e] += float(tr) + xs[x] += float(tr) + return (es, xs) diff --git a/tools/gpu_backtest_tools/runpod/__init__.py b/tools/gpu_backtest_tools/runpod/__init__.py new file mode 100644 index 0000000..23986ca --- /dev/null +++ b/tools/gpu_backtest_tools/runpod/__init__.py @@ -0,0 +1,13 @@ +"""Optional RunPod helper; the core never imports allocation or transport.""" + +from .client import RunPodClient, RunPodError, load_api_key +from .launcher import DEFAULT_IMAGE, launch, terminate_on_signal + +__all__ = [ + "DEFAULT_IMAGE", + "RunPodClient", + "RunPodError", + "launch", + "load_api_key", + "terminate_on_signal", +] diff --git a/tools/gpu_backtest_tools/runpod/bundle.py b/tools/gpu_backtest_tools/runpod/bundle.py new file mode 100644 index 0000000..bd36eae --- /dev/null +++ b/tools/gpu_backtest_tools/runpod/bundle.py @@ -0,0 +1,215 @@ +"""Select upload files and construct the remote execution recipe.""" + +import hashlib +import importlib.metadata +import importlib.resources +import json +import shutil +import tarfile +import tempfile +from pathlib import Path + +from gpu_backtest import __version__, resolve_strategy_name + +PACKAGE_FILES = { + "gpu_backtest": [ + "__init__.py", + "__main__.py", + "cli/__init__.py", + "cli/analysis.py", + "cli/backtest.py", + "cli/options.py", + "cli/tools.py", + "core/__init__.py", + "core/data.py", + "core/engine.py", + "core/grid.py", + "core/indicators.py", + "core/kernels.py", + "core/output.py", + "core/statistics.py", + "core/strategy.py", + "workflows/__init__.py", + "workflows/charts.py", + "workflows/common.py", + "workflows/pipeline.py", + "workflows/splits.py", + ], + "gpu_backtest_examples": [ + "__init__.py", + "data.py", + "rsi/__init__.py", + "rsi/generate_data.py", + "rsi/strategy.py", + ], + "gpu_backtest_tools": [ + "__init__.py", + "benchmarks/__init__.py", + "benchmarks/cpu.py", + "benchmarks/runner.py", + "checks/__init__.py", + "checks/gpu.py", + "checks/reference.py", + "runpod/__init__.py", + "runpod/bundle.py", + "runpod/client.py", + "runpod/launcher.py", + "runpod/transport.py", + "runpod/requirements.txt", + ], +} +CONFIG_KEYS = { + "num_splits", + "exit_dims", + "input", + "buy", + "strategy", + "top_n", + "end_date", + "entry_dims", + "start_date", + "expected_interval", + "threads_per_block", + "ranges", + "sell", +} +IGNORED_PARTS = {"__pycache__", ".git", ".ruff_cache", "venv", ".venv", ".pytest_cache"} + + +def build_bundle(archive, *, mode, config_path=None, plugin_dir=None, charts=False): + if mode not in ("run", "pipeline", "check", "benchmark"): + raise ValueError("RunPod mode must be run, pipeline, check, or benchmark") + config = {} + if mode in ("run", "pipeline"): + if not config_path: + raise ValueError("RunPod run/pipeline requires --config") + config_path = Path(config_path).resolve() + config = json.loads(config_path.read_text()) + if not isinstance(config, dict) or set(config) - CONFIG_KEYS: + raise ValueError("RunPod config must contain only documented engine config keys") + if not config.get("strategy") or not config.get("input"): + raise ValueError("RunPod config must specify strategy and input") + if not isinstance(config["strategy"], str) or not all( + (p.isidentifier() for p in config["strategy"].split(".")) + ): + raise ValueError("RunPod strategy must be a builtin name or dotted module name") + source = (config_path.parent / config["input"]).resolve() + from gpu_backtest.core.data import validate_market_data + + validate_market_data(source, expected_interval=config.get("expected_interval")) + distribution = importlib.metadata.distribution("gpu-backtest-engine") + base_requires = [r for r in distribution.requires or [] if "extra ==" not in r] + with tempfile.TemporaryDirectory(prefix="runpod-bundle-") as directory: + staging = Path(directory) + engine = staging / "engine" + for package_name, members in PACKAGE_FILES.items(): + package = Path(str(importlib.resources.files(package_name))) + for relative in members: + path = package / relative + if path.is_symlink(): + raise ValueError("Engine bundle cannot include symlinked source files") + output = engine / package_name / relative + output.parent.mkdir(parents=True, exist_ok=True) + shutil.copyfile(path, output) + project = f"""[build-system] +requires = ["setuptools>=77", "wheel"] +build-backend = "setuptools.build_meta" + +[project] +name = "gpu-backtest-engine" +version = {json.dumps(__version__)} +requires-python = ">=3.11,<3.14" +dependencies = {json.dumps(base_requires)} + +[project.scripts] +gpu-backtest = "gpu_backtest.cli:main" + +[tool.setuptools.packages.find] +where = ["."] + +[tool.setuptools.package-data] +"gpu_backtest_tools.runpod" = ["requirements.txt"] +""" + (engine / "pyproject.toml").write_text(project) + for entry in distribution.files or []: + if str(entry) == "LICENSE" or str(entry).endswith("/licenses/LICENSE"): + shutil.copyfile(distribution.locate_file(entry), engine / "LICENSE") + break + if mode in ("run", "pipeline"): + (staging / "input").mkdir() + shutil.copyfile(source, staging / "input/market.csv") + config["input"] = "input/market.csv" + (staging / "job.json").write_text(json.dumps(config, indent=2) + "\n") + if plugin_dir: + plugin_dir = Path(plugin_dir).resolve() + if not plugin_dir.is_dir(): + raise ValueError("Plugin directory does not exist") + copied = 0 + for path in sorted(plugin_dir.rglob("*.py")): + relative = path.relative_to(plugin_dir) + if set(relative.parts) & IGNORED_PARTS: + continue + if path.is_symlink(): + raise ValueError("Plugin source symlinks are not supported") + output = staging / "plugins" / relative + output.parent.mkdir(parents=True, exist_ok=True) + shutil.copyfile(path, output) + copied += 1 + if not copied: + raise ValueError("Plugin directory contains no Python files") + if mode in ("run", "pipeline"): + config["strategy"] = resolve_strategy_name(config["strategy"]) + (staging / "job.json").write_text(json.dumps(config, indent=2) + "\n") + if ( + mode in ("run", "pipeline") + and config["strategy"] != "gpu_backtest_examples.rsi.strategy" + ): + relative = Path(*config["strategy"].split(".")) + if not (staging / "plugins" / relative.with_suffix(".py")).is_file() and ( + not (staging / "plugins" / relative / "__init__.py").is_file() + ): + raise ValueError("External strategy must be present in --plugin-dir") + (staging / "job.sh").write_text(remote_script(mode, charts=charts)) + files = [] + for path in sorted(staging.rglob("*")): + if path.is_file(): + files.append( + { + "file": path.relative_to(staging).as_posix(), + "sha256": hashlib.sha256(path.read_bytes()).hexdigest(), + } + ) + (staging / "bundle_manifest.json").write_text(json.dumps(files, indent=2) + "\n") + with tarfile.open(archive, "w:gz") as handle: + for path in sorted(staging.rglob("*")): + if path.is_file(): + handle.add(path, arcname=path.relative_to(staging).as_posix(), recursive=False) + return files + + +def remote_script(mode, *, charts=False): + script = """#!/bin/bash +set -euo pipefail +export NUMBA_ENABLE_CUDASIM=0 +export PYTHONPATH="$PWD/plugins${PYTHONPATH:+:$PYTHONPATH}" +mkdir -p output +python3 -c 'import sys; assert (3,11) <= sys.version_info[:2] < (3,14), "Python 3.11-3.13 required"' +python3 -m venv env +env/bin/python -m pip install --disable-pip-version-check ./engine -r engine/gpu_backtest_tools/runpod/requirements.txt +nvidia-smi --query-gpu=name,driver_version --format=csv > output/gpu.csv +env/bin/python -m gpu_backtest gpu-check --output output/gpu_check.json +""" + if mode == "run": + script += ( + "env/bin/python -m gpu_backtest run --config job.json --out-prefix output/result\n" + ) + elif mode == "pipeline": + script += "env/bin/python -m gpu_backtest pipeline --config job.json --output-dir output/pipeline\n" + elif mode == "benchmark": + script += "env/bin/python -m gpu_backtest benchmark --output output/benchmark.json --billion --billion-cpu\n" + if charts and mode in ("run", "pipeline"): + folder = "output/pipeline" if mode == "pipeline" else "output" + script += "env/bin/python -m pip install 'altair==6.3.0'\n" + script += f"env/bin/python -m gpu_backtest charts --input {folder}/*_top_entry.csv {folder}/*_top_exit.csv --output output/charts.html\n" + script += "env/bin/python -m pip freeze > output/requirements-resolved.txt\n" + return script diff --git a/tools/gpu_backtest_tools/runpod/client.py b/tools/gpu_backtest_tools/runpod/client.py new file mode 100644 index 0000000..3ef9233 --- /dev/null +++ b/tools/gpu_backtest_tools/runpod/client.py @@ -0,0 +1,99 @@ +"""RunPod authentication and idempotent API operations.""" + +import json +import os +import time +import tomllib +import urllib.error +import urllib.request +from pathlib import Path + +from gpu_backtest import __version__ + +API_URL = "https://rest.runpod.io/v1" + + +class RunPodError(RuntimeError): + pass + + +def load_api_key(): + key = os.environ.get("RUNPOD_API_KEY") + if not key: + path = Path.home() / ".runpod/config.toml" + if path.is_file(): + key = tomllib.loads(path.read_text()).get("apikey") + if not isinstance(key, str) or not key.strip(): + raise ValueError("Set RUNPOD_API_KEY or apikey in ~/.runpod/config.toml") + return key.strip() + + +class RunPodClient: + def __init__(self, key): + self.key = key + + def request(self, method, path, body=None): + request = urllib.request.Request( + API_URL + path, + data=json.dumps(body).encode() if body is not None else None, + headers={ + "Authorization": f"Bearer {self.key}", + "Content-Type": "application/json", + "Accept": "application/json", + "User-Agent": f"gpu-backtest-engine/{__version__}", + }, + method=method, + ) + try: + with urllib.request.urlopen(request, timeout=30) as response: + payload = response.read() + return json.loads(payload) if payload else None + except urllib.error.HTTPError as exc: + if method == "DELETE" and exc.code == 404: + return None + raise RunPodError(f"RunPod {method} {path}: HTTP {exc.code}") from None + except (urllib.error.URLError, TimeoutError, json.JSONDecodeError) as exc: + raise RunPodError(f"RunPod {method} {path}: connection or response failure") from exc + + def create(self, body): + try: + result = self.request("POST", "/pods", body) + if ( + not isinstance(result, dict) + or not isinstance(result.get("id"), str) + or (not result["id"].isascii()) + or (not result["id"].isalnum()) + ): + raise RunPodError("Creation response did not contain a valid pod ID") + except RunPodError as exc: + if "HTTP 4" in str(exc): + raise + try: + pods = self.request("GET", "/pods") + matches = [p for p in pods if p.get("name") == body["name"]] + except (RunPodError, TypeError): + matches = [] + if ( + len(matches) == 1 + and isinstance(matches[0].get("id"), str) + and matches[0]["id"].isascii() + and matches[0]["id"].isalnum() + ): + return matches[0] + raise RunPodError( + f"Creation outcome unknown; check RunPod for launch {body['name']}. No second creation request was sent." + ) from exc + return result + + def get(self, pod_id): + return self.request("GET", "/pods/" + pod_id) + + def delete(self, pod_id): + for attempt in range(3): + try: + self.request("DELETE", "/pods/" + pod_id) + return + except RunPodError: + if attempt == 2: + raise + time.sleep(attempt + 1) diff --git a/tools/gpu_backtest_tools/runpod/launcher.py b/tools/gpu_backtest_tools/runpod/launcher.py new file mode 100644 index 0000000..f42cf12 --- /dev/null +++ b/tools/gpu_backtest_tools/runpod/launcher.py @@ -0,0 +1,229 @@ +"""Own the paid-pod lifecycle, deadlines, cancellation, and cleanup state.""" + +import contextlib +import json +import math +import shutil +import signal +import subprocess +import threading +import time +import uuid +from pathlib import Path + +from .bundle import build_bundle +from .client import RunPodClient, RunPodError, load_api_key +from .transport import SSHTransport, endpoint + +DEFAULT_IMAGE = "runpod/pytorch:2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04" + + +def _remaining(deadline): + seconds = deadline - time.monotonic() + if seconds <= 0: + raise RunPodError("RunPod job exceeded its local time limit") + return seconds + + +def _write_state(path, **fields): + state = json.loads(path.read_text()) if path.exists() else {} + state.update(fields) + temporary = path.with_suffix(".tmp") + temporary.write_text(json.dumps(state, indent=2) + "\n") + temporary.replace(path) + + +def _record_cleanup_state(path, **fields): + """Failure to write metadata must not prevent or misreport API cleanup.""" + try: + _write_state(path, **fields) + except (OSError, ValueError) as exc: + print(f"Could not update local state {path}: {type(exc).__name__}", flush=True) + + +@contextlib.contextmanager +def defer_creation_signals(): + """Record Ctrl-C/termination until ownership of a creation reply is recorded.""" + interrupted = [False] + if threading.current_thread() is not threading.main_thread(): + yield interrupted + return + previous = {number: signal.getsignal(number) for number in (signal.SIGINT, signal.SIGTERM)} + + def defer(signum, frame): + interrupted[0] = True + + for number in previous: + signal.signal(number, defer) + try: + yield interrupted + finally: + for number, handler in previous.items(): + signal.signal(number, handler) + + +@contextlib.contextmanager +def terminate_on_signal(): + previous = signal.getsignal(signal.SIGTERM) + + def interrupted(signum, frame): + raise KeyboardInterrupt("RunPod launcher interrupted") + + signal.signal(signal.SIGTERM, interrupted) + try: + yield + finally: + signal.signal(signal.SIGTERM, previous) + + +def launch( + *, + config_path=None, + output_dir, + ssh_key=None, + mode="pipeline", + plugin_dir=None, + image=DEFAULT_IMAGE, + gpu="NVIDIA GeForce RTX 4090", + cloud="SECURE", + max_seconds=1800, + max_hourly_rate=1.0, + keep_pod=False, + charts=False, + dry_run=False, + client=None, + transport_factory=SSHTransport, +): + if not isinstance(max_seconds, int) or max_seconds < 1: + raise ValueError("max_seconds must be a positive integer") + if not math.isfinite(max_hourly_rate) or max_hourly_rate <= 0: + raise ValueError("max_hourly_rate must be finite and positive") + if cloud not in ("SECURE", "COMMUNITY"): + raise ValueError("cloud must be SECURE or COMMUNITY") + output_dir = Path(output_dir).resolve() + if output_dir.exists() and any(output_dir.iterdir()): + raise ValueError("RunPod output directory must be new or empty") + output_dir.mkdir(parents=True, exist_ok=True) + files = build_bundle( + output_dir / "upload.tar.gz", + mode=mode, + config_path=config_path, + plugin_dir=plugin_dir, + charts=charts, + ) + if dry_run: + print( + json.dumps( + { + "image": image, + "gpu": gpu, + "cloud": cloud, + "files": [f["file"] for f in files], + "archive": str(output_dir / "upload.tar.gz"), + "pod_created": False, + }, + indent=2, + ) + ) + return output_dir + if not ssh_key or not Path(ssh_key).expanduser().is_file(): + raise ValueError("Specify an existing --ssh-key; add its public key to RunPod Credentials") + key = Path(ssh_key).expanduser().resolve() + if not shutil.which("ssh"): + raise ValueError("An OpenSSH client is required") + client = client or RunPodClient(load_api_key()) + launch_name = "gpu-backtest-" + uuid.uuid4().hex[:12] + body = { + "name": launch_name, + "imageName": image, + "gpuTypeIds": [gpu], + "gpuCount": 1, + "cloudType": cloud, + "computeType": "GPU", + "containerDiskInGb": 30, + "volumeInGb": 0, + "ports": ["22/tcp"], + "supportPublicIp": True, + } + if mode == "benchmark": + body["minVCPUPerGPU"] = 8 + public_key = Path(str(key) + ".pub") + if public_key.is_file(): + body["env"] = {"PUBLIC_KEY": " ".join(public_key.read_text().split()[:2])} + state_path = output_dir / "runpod_state.json" + _write_state(state_path, name=launch_name, image=image, gpu=gpu, status="creating") + pod_id = None + failure = None + keep_allowed = False + print(f"Creating one {gpu} pod on {cloud}...", flush=True) + try: + with defer_creation_signals() as cancelled: + pod = client.create(body) + candidate_id = pod["id"] + if not isinstance(candidate_id, str) or not candidate_id.isalnum(): + raise RunPodError("RunPod returned an invalid pod ID") + pod_id = candidate_id + _write_state(state_path, pod_id=pod_id, status="starting") + if cancelled[0]: + raise KeyboardInterrupt("RunPod creation was cancelled") + deadline = time.monotonic() + max_seconds + quoted_rate = pod.get("adjustedCostPerHr") + if quoted_rate is None: + quoted_rate = pod.get("costPerHr") + rate = float(quoted_rate) if quoted_rate is not None else math.inf + if not math.isfinite(rate) or rate < 0 or rate > max_hourly_rate: + raise RunPodError(f"Pod rate exceeds the requested {max_hourly_rate:g}/hour limit") + keep_allowed = True + _write_state(state_path, hourly_rate=rate) + print(f"Pod {pod_id}: ${rate:g}/hour. Waiting for SSH...", flush=True) + while True: + _remaining(deadline) + details = client.get(pod_id) + connection = endpoint(details) + if connection: + transport = transport_factory(*connection, key, output_dir) + try: + if transport.ready(): + break + except subprocess.TimeoutExpired: + pass + time.sleep(min(5, _remaining(deadline))) + workspace = "/workspace/" + launch_name + _write_state(state_path, status="uploading") + print("Uploading selected engine/data/plugin files...", flush=True) + transport.upload(output_dir / "upload.tar.gz", workspace, output_dir, _remaining(deadline)) + _write_state(state_path, status="running") + print( + "Installing environment and running GPU checks/job; progress is in remote.log...", + flush=True, + ) + transport.execute(workspace, output_dir, _remaining(deadline)) + _write_state(state_path, status="downloading") + print("Downloading results...", flush=True) + transport.download(workspace, output_dir, _remaining(deadline)) + _write_state(state_path, status="completed") + except BaseException as exc: + failure = exc + _record_cleanup_state(state_path, status="failed", error=type(exc).__name__) + raise + finally: + if pod_id: + if keep_pod and keep_allowed: + _record_cleanup_state(state_path, cleanup="kept") + print( + f"Pod {pod_id} retained by --keep-pod; it continues to incur charges.", + flush=True, + ) + else: + try: + client.delete(pod_id) + except BaseException as exc: + _record_cleanup_state(state_path, cleanup="failed") + message = f"Could not delete pod {pod_id}; delete it in RunPod Console. State: {state_path}" + print(message, flush=True) + if failure is None: + raise RunPodError(message) from exc + else: + _record_cleanup_state(state_path, cleanup="deleted") + print(f"Pod {pod_id} deleted.", flush=True) + return output_dir diff --git a/src/gpu_backtest/runpod_requirements.txt b/tools/gpu_backtest_tools/runpod/requirements.txt similarity index 100% rename from src/gpu_backtest/runpod_requirements.txt rename to tools/gpu_backtest_tools/runpod/requirements.txt diff --git a/tools/gpu_backtest_tools/runpod/transport.py b/tools/gpu_backtest_tools/runpod/transport.py new file mode 100644 index 0000000..78d6783 --- /dev/null +++ b/tools/gpu_backtest_tools/runpod/transport.py @@ -0,0 +1,140 @@ +"""SSH transport and safe result extraction; no allocation logic.""" + +import contextlib +import ipaddress +import json +import os +import shlex +import shutil +import subprocess +import tarfile +from pathlib import Path + +from .client import RunPodError + + +def endpoint(pod): + address = pod.get("publicIp") + port = (pod.get("portMappings") or {}).get("22") + if not address or not port: + return None + try: + address = str(ipaddress.ip_address(address)) + port = int(port) + if not 1 <= port <= 65535: + raise ValueError + except (ValueError, TypeError): + raise RunPodError("RunPod returned invalid public SSH connection details") from None + return (address, port) + + +class SSHTransport: + def __init__(self, address, port, key, output_dir): + self.base = [ + "ssh", + "-i", + str(key), + "-p", + str(port), + "-o", + "BatchMode=yes", + "-o", + "IdentitiesOnly=yes", + "-o", + "ForwardAgent=no", + "-o", + "ClearAllForwardings=yes", + "-o", + "StrictHostKeyChecking=accept-new", + "-o", + f"UserKnownHostsFile={json.dumps(str(Path(output_dir) / 'known_hosts'))}", + "-o", + "ConnectTimeout=10", + "-o", + "ServerAliveInterval=15", + "-o", + "ServerAliveCountMax=3", + f"root@{address}", + ] + self.environment = os.environ.copy() + self.environment.pop("RUNPOD_API_KEY", None) + + def ready(self): + result = subprocess.run( + [*self.base, "true"], + stdin=subprocess.DEVNULL, + capture_output=True, + timeout=15, + env=self.environment, + ) + return result.returncode == 0 + + def _run(self, command, *, stdin=None, stdout=None, log=None, timeout): + with contextlib.ExitStack() as stack: + input_file = ( + stack.enter_context(Path(stdin).open("rb")) if stdin else subprocess.DEVNULL + ) + output_file = stack.enter_context(Path(stdout).open("wb")) if stdout else None + log_file = stack.enter_context(Path(log).open("ab")) if log else None + process = subprocess.Popen( + [*self.base, command], + stdin=input_file, + stdout=output_file or log_file, + stderr=log_file, + env=self.environment, + ) + try: + code = process.wait(timeout=timeout) + except BaseException: + process.kill() + process.wait() + raise + if code: + raise RunPodError(f"SSH stage failed with exit code {code}; inspect remote.log") + + def upload(self, archive, workspace, output_dir, timeout): + command = f"mkdir -p {shlex.quote(workspace)} && tar xzf - --no-same-owner -C {shlex.quote(workspace)}" + self._run(command, stdin=archive, log=Path(output_dir) / "remote.log", timeout=timeout) + + def execute(self, workspace, output_dir, timeout): + self._run( + f"cd {shlex.quote(workspace)} && bash job.sh", + log=Path(output_dir) / "remote.log", + timeout=timeout, + ) + + def download(self, workspace, output_dir, timeout): + archive = Path(output_dir) / "results.tar.gz" + self._run( + f"cd {shlex.quote(workspace)}/output && tar czf - .", + stdout=archive, + log=Path(output_dir) / "remote.log", + timeout=timeout, + ) + extract_results(archive, Path(output_dir) / "results") + + +def extract_results(archive, destination): + """Accept ordinary result files only; remote tar metadata is untrusted.""" + destination = Path(destination) + destination.mkdir(parents=True, exist_ok=True) + with tarfile.open(archive, "r:gz") as handle: + members = handle.getmembers() + for member in members: + path = Path(member.name) + if ( + path.is_absolute() + or ".." in path.parts + or (not (member.isfile() or member.isdir())) + ): + raise RunPodError("Downloaded results contain an unsafe tar entry") + if not (destination / path).resolve().is_relative_to(destination.resolve()): + raise RunPodError("Downloaded results escape their destination") + for member in members: + path = destination / member.name + if member.isdir(): + path.mkdir(parents=True, exist_ok=True) + else: + path.parent.mkdir(parents=True, exist_ok=True) + with path.open("wb") as output: + shutil.copyfileobj(handle.extractfile(member), output)