Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 14 additions & 8 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,18 @@
# Contributing

Install `.[dev,viz]`, then run `pytest`, `ruff check .`, and `ruff format --check .`.
Keep simulator tests in a separate process from real-GPU checks. For numerical
changes, include an explicit expected result or independent reference.
Install `.[dev,viz]`; run `pytest`, `ruff check .`, and `ruff format --check .`.
See [tests/README.md](tests/README.md) for CPU, simulator, and hardware commands.

New example strategies need a documented device-function spec and hand-calculated
trade case. Keep demonstrations small and use synthetic data. Do not include
credentials, personal market datasets, private strategy plugins, or research output.
Keep boundaries clear:

Document numerical and execution-rule changes in the same pull request. New tracked
files must be listed in `scripts/check_release.py` after reviewing their contents.
- `src/gpu_backtest/core/` contains generic computation only. It must not import
examples, test references, benchmark code, CLI, workflows, or deployment tools.
- `workflows/` uses core primitives for splitting/analysis/output charts.
- Example trading rules/data stay under `examples/` and helper code under `tools/`.
- CLI modules are thin adapters; they do not own numerical or cloud logic.

New example strategies need a device-function spec and hand-calculated case.
Use synthetic data. Keep credentials, private plugins/datasets, and research output
out of Git. New tracked files require a reviewed update to the explicit allowlist
in `scripts/check_release.py`. Preserve historical benchmark JSONs; new claims need
new measured evidence. Document intentional numerical/contract changes in the PR.
8 changes: 4 additions & 4 deletions MANIFEST.in
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
include LICENSE README.md CONTRIBUTING.md
recursive-include src/gpu_backtest *.py
recursive-include src/gpu_backtest *.txt
recursive-include tests *.py
recursive-include tools/gpu_backtest_tools *.py *.txt
recursive-include examples/gpu_backtest_examples *.py *.json *.csv *.md
recursive-include tests *.py *.md
recursive-include docs *.md
recursive-include benchmarks *.json
recursive-include examples *.py *.json *.csv
recursive-include benchmarks *.json *.md
recursive-include scripts *.py
global-exclude __pycache__ *.py[cod] .env *.pem *.key
242 changes: 75 additions & 167 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,131 +2,64 @@

## 10× faster on our public billion-pair benchmark

**CPU: 7 min 44.60 s → RTX 4090: 44.95 s.** Measured on the same
**1,000,000,000-pair RSI grid × 1,024 bars**, against an **eight-thread compiled
Numba CPU baseline**. Exact measured speedup: **10.34×**; cloud setup is additional.
**CPU: 7 min 44.60 s → RTX 4090: 44.95 s.** Same RSI grid:
**1,000,000,000 pairs × 1,024 bars**, compared with an **eight-thread compiled
Numba CPU baseline**. Measured speedup **10.34×**, saving about seven minutes per
sweep. Cloud setup is additional; [method and raw evidence](docs/benchmarks.md).

**Backtest your own trading algorithms. Sweep a billion parameter combinations on a GPU.**
- **Your algorithm:** load a strategy module/object; private rules can remain private.
- **Large grids:** deterministic GPU reductions without materializing the full return matrix.
- **Your machine or RunPod:** use a local NVIDIA GPU or the separate cloud helper.

- **Swap algorithms:** load a separately installed strategy module or object.
Your strategy can stay in a private repo; the engine does not need to be edited.
- **10× faster — CPU minutes → GPU seconds:** the same **billion-pair RSI sweep** took
**7 min 44.60 s on an eight-thread Numba CPU baseline → 44.95 s on RTX 4090**.
That's **10.34× faster**, saving approximately **seven minutes per sweep**.
- **Billion-scale sweep:** **1,000,000,000 unique pairs × 1,024 bars**, with
both sides measured through statistics and CSV/manifest output.
- **No GPU in your laptop:** use the RunPod launcher to rent, run, download, and clean up.
## Where everything lives

This is a measured full-grid comparison, not an extrapolation from a tiny case.
See [the benchmark and raw data](docs/benchmarks.md) for hardware, timing scope,
and reproduction. Results depend on workload; cloud setup time is additional.

The engine provides deterministic reductions and separate entry/exit effect-size
rankings without storing a full pairwise return matrix. This distribution includes
one educational RSI strategy and generated synthetic OHLCV data. Custom plugins
must follow [the Numba device-function contract](docs/strategy-contract.md).

## Published performance

| Public RSI workload | Compiled CPU, eight threads | RTX 4090 | Time saved |
|---|---|---|---|
| **1,000,000,000 pairs × 1,024 bars** | **7 min 44.60 s** | **44.95 s** | **6 min 59.65 s per sweep; 10.34× faster** |

The CPU is an AMD EPYC 7K62 host running a parallel, compiled Numba baseline.
Both measurements use the same data, grid, fees, and two-pass algorithm, through
ranked CSV/manifest output. The CPU reducer and CUDA context were already warmed;
GPU kernel construction/JIT is included. Each full-grid timing is one measured run.
These are engine-job times, excluding pod provisioning and installation.

**When GPU is useful:** repeated large parameter searches, where saving minutes
on every sweep adds up. If your job already finishes in a few CPU seconds,
renting/setup overhead may outweigh GPU savings. The smaller warmed-kernel timing
tests remain in the detailed benchmark, rather than being the main use-case claim.
```text
src/gpu_backtest/
core/ GPU engine, kernels, grids, indicators, statistics, output
workflows/ Generic split/common analysis and charts
cli/ Thin command-line adapters
examples/gpu_backtest_examples/
rsi/ Example strategy + config + generated CSV
data.py Example/benchmark synthetic data generator
tools/gpu_backtest_tools/
runpod/ Optional API / SSH / bundle / lifecycle helper
benchmarks/ Performance runner and compiled CPU baseline
checks/ Hardware smoke checks and CPU test reference
tests/
cpu/ CPU contracts, examples, workflows, helper tests
gpu/ Isolated CUDA simulation and real-GPU tests
docs/ Strategy contract, RunPod usage, benchmark method
benchmarks/results/ Historical measurements and validation evidence
```

The billion grid uses 20,000 entry sets × 50,000 exit sets. Its grouped arrays
occupy 1.12 MB, compared with 4 GB for a full float32 return matrix; this excludes
input/indicator tables and runtime overhead. Pair counts are parameter combinations,
not trades. These measurements use public code and synthetic data, with no private
algorithm or market dataset.
**Start with `core/engine.py`** for the GPU run. Numerical kernels are in
`core/kernels.py`; trading rules are supplied by a plugin. Core imports no example,
CPU comparison engine, benchmark, or RunPod code. CPU preprocessing of market data
and indicator tables is part of the GPU pipeline; the separate CPU backtest baseline
is only a benchmark/reference tool.

## Quickstart without a GPU
## Try the RSI example without a GPU

Python 3.11–3.13 is supported. From a checkout:
Python 3.11–3.13, from a checkout:

```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev,viz]'
NUMBA_ENABLE_CUDASIM=1 gpu-backtest run \
--config examples/rsi.json --out-prefix runs/rsi
--config examples/gpu_backtest_examples/rsi/config.json --out-prefix runs/rsi
gpu-backtest charts --input runs/rsi_top_entry.csv runs/rsi_top_exit.csv \
--output runs/rsi.html
```

Open `runs/rsi.html` in a browser. Its Vega libraries load from a public CDN.
The CUDA simulator is for tiny demonstrations and correctness tests. Large grids
need a real NVIDIA GPU.

The example sweeps four entry combinations against four exit combinations.
It writes two ranked CSVs and an output manifest:

```text
runs/rsi_top_entry.csv
runs/rsi_top_exit.csv
runs/rsi_top_manifest.json
```

`examples/synthetic.csv` contains generated bars, not historical market data.
Regenerate it with `python examples/generate_data.py`.

## RunPod: use a GPU from your laptop

The same engine runs on a rented RunPod GPU. With a funded account, Pod API key,
and registered SSH key:

```bash
gpu-backtest runpod --config examples/rsi.json \
--ssh-key ~/.ssh/runpod_ed25519 --output-dir runs/runpod-rsi --charts
```

This leases one RTX 4090, uploads selected engine/data/plugin files, installs the
environment, runs hardware numeric checks and the pipeline, downloads output,
and deletes the pod. Add `--dry-run` to inspect the bundle without renting anything;
`--mode run` selects one sweep and `--mode check` runs only hardware checks.
Private modules can be supplied with `--plugin-dir` without adding them to this repo.
See [the RunPod guide](docs/runpod.md) for setup, the manual route, environment
requirements, time/rate limits, custom plugins, and cleanup behavior.

## NVIDIA GPU setup (inside the GPU machine)

Run these commands in the GPU computer's terminal. For RunPod, use the pinned
setup in [the RunPod guide](docs/runpod.md); your laptop only controls the launcher.
For another GPU workstation/server, install a compatible driver and CUDA runtime/toolkit. The separate NVIDIA
`numba-cuda` backend uses the same `from numba import cuda` interface:

```bash
python -m pip install -e '.[cuda]'
# If you also need CUDA 12 Python runtime/toolkit dependencies:
python -m pip install 'numba-cuda[cu12]>=0.30,<0.31'
gpu-backtest run --config examples/rsi.json --out-prefix runs/rsi_gpu
```

See the [NVIDIA installation guide](https://nvidia.github.io/numba-cuda/user/installation.html)
for CUDA 12/13 and driver requirements. Run GPU commands without
`NUMBA_ENABLE_CUDASIM=1`. The project's tested Numba series is 0.65; upgrading the
backend requires rerunning the simulator and real-GPU checks.
The simulator is for tiny examples/tests. Open `runs/rsi.html` in a browser;
its Vega libraries load from a public CDN. The [RSI example](examples/gpu_backtest_examples/rsi/README.md)
is educational and uses generated data. It is packaged separately as
`gpu_backtest_examples.rsi.strategy`, not as an engine builtin.

## Use your own strategy

Install the package containing your strategy in the same Python environment:

```bash
python -m pip install -e /path/to/your-strategy-package
gpu-backtest run --strategy my_strategies.example \
--input /path/to/market.csv --out-prefix runs/custom
```

Or pass a module/object through the Python API:
Install your strategy package into the same environment:

```python
from gpu_backtest import run
Expand All @@ -135,55 +68,44 @@ from my_strategies import example
results = run(example, "market.csv", "runs/custom", buy=0.0015, sell=0.0015)
```

The engine loads an installed Python module. Plugins are trusted code, and run
locally; they must implement the documented Numba device-function contract.
See [the strategy contract](docs/strategy-contract.md) for parameter specifications,
indicator tables, CPU references, and the RSI execution rules.
Or use `gpu-backtest run --strategy my_strategies.example --input market.csv
--out-prefix runs/custom`. Plugins implement a complete Numba CUDA trading loop;
see [the contract](docs/strategy.md). Entry/exit parameters are separate Cartesian
axes (up to four dimensions each), with up to four precomputed tables. Shared-parameter
diagonal-only sweeps are not supported. The API returns grouped sums/squares and
optional ranked artifacts, not a trade ledger, equity curve, Sharpe, or drawdown series.

## Optional workflows and helpers

## Date splits and common parameters
| Command | Owner / purpose |
|---|---|
| `pipeline` | `workflows/`: split, run each segment, intersect ranked parameters |
| `split`, `common`, `charts` | `workflows/`: standalone data/result analysis |
| `runpod` | `tools/runpod/`: lease one GPU, upload selected files, download, delete |
| `benchmark` | `tools/benchmarks/`: reproduce the public CPU/GPU measurements |
| `gpu-check` | `tools/checks/`: small numeric checks on actual hardware |

```bash
NUMBA_ENABLE_CUDASIM=1 gpu-backtest pipeline \
--config examples/rsi.json --output-dir runs/pipeline
gpu-backtest runpod --config examples/gpu_backtest_examples/rsi/config.json \
--ssh-key ~/.ssh/runpod_ed25519 --output-dir runs/runpod-rsi --charts
```

With three splits, the pipeline evaluates the full period and its two disjoint
halves. It records split hashes, writes ranked artifacts for every split, and
finds parameter combinations appearing in every top artifact. The output includes
`common_entry.csv` and `common_exit.csv` with per-period, average, and minimum
effect sizes. No common combinations is reported as a failed analysis, not an
empty successful result.

Config `input` paths resolve relative to the JSON file. CLI `--input` and output
paths resolve relative to the current directory. A strategy must always be
specified explicitly. Config keys are `strategy`, `input`, `buy`, `sell`,
`entry_dims`, `exit_dims`, `top_n`, `threads_per_block`, `expected_interval`,
`num_splits`, `start_date`, `end_date`, and `ranges`. Dates use `DD-MM-YYYY`;
custom ranges are arrays such as `[["01-01-2024", "31-01-2024"], ...]`.

`gpu-backtest split`, `common`, and `charts` also work independently; run each
subcommand with `--help` for its arguments.

## Scope and interpretation

- Entry and exit parameters form a Cartesian product, with at most four
dimensions on each side and four precomputed indicator tables. Shared-parameter
diagonal-only sweeps are not supported.
- The engine runs two passes over the same return matrix. Returns and squared
returns are rounded to float32, then accumulated in float64 in a fixed order.
- Ranked rows represent entry/exit parameter groups, not individually selected
complete strategies. Their observations are parameter combinations, not
independent market samples. Effect sizes and nominal normal intervals are
descriptive grid comparisons. `posterior_prob_superior` is a normal-CDF proxy;
it is not a Bayesian probability or a forecast of profitable trading.
- Strategy plugins own their trading loop, position sizing, and execution rules.
The included example uses long-only next-open execution, with its boundary and
fee conventions documented in the strategy contract.
- The current API returns grouped sums and sums of squares, plus optional top
artifacts. It does not return a trade ledger, equity curve, Sharpe ratio, or
drawdown series. Example performance is not a trading recommendation.

## Development
RunPod requires your account/API key and registered SSH key. It is an optional
execution helper; [setup, manual GPU route, and cleanup](docs/runpod.md).
Commands keep their existing names. The public `from gpu_backtest import run` API
and old `rsi_meanrev` shorthand remain usable; direct internal imports moved under
`core/`, `workflows/`, or `gpu_backtest_tools` in v0.5.

## Performance and development

| Same billion-pair RSI job | CPU, eight threads | RTX 4090 | Saved per sweep |
|---|---|---|---|
| 1,000,000,000 pairs × 1,024 bars, through output | 7 min 44.60 s | 44.95 s | 6 min 59.65 s; 10.34× faster |

This is one measured run per side on the published setup, not a universal speed
claim. Small CPU jobs may not justify cloud startup. Statistics describe parameter
combination groups, not independent market samples or a forecast of profits.
See [benchmark details](docs/benchmarks.md) for scope and raw data.

```bash
python -m pytest
Expand All @@ -193,21 +115,7 @@ python -m build
python scripts/check_release.py
```

Default tests need no GPU. They run simulation in a separate process and verify
hand-calculated results, known reduction matrices, independent CPU references,
output determinism, statistics, indicator values, split coverage, and external
plugin loading. For a real NVIDIA GPU:

```bash
NUMBA_ENABLE_CUDASIM=0 python -m pytest -m gpu -v
```

See [testing details](docs/testing.md) and the [RunPod validation record](docs/runpod-validation.md).
Small numeric checks and the public RSI pipeline have passed on a real RTX 4090.
The [public benchmark](docs/benchmarks.md) also completed a billion-pair RSI sweep.
Other hardware/backend versions and custom strategies require their own validation.

## License
[CPU/GPU testing](tests/README.md) · [Plugin contract](docs/strategy.md) ·
[RunPod helper](docs/runpod.md) · [Benchmark method](docs/benchmarks.md)

[MIT](LICENSE). You may use the engine with separately maintained private
strategy plugins, subject to their own licenses.
[MIT](LICENSE). Private strategy plugins retain their own licenses.
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ confirmed it was absent from the account's pod list.
| Optional HTML renderer | Altair 6.3.0 |

The launcher now pins the resolved core packages in
[`runpod_requirements.txt`](../src/gpu_backtest/runpod_requirements.txt).
[`runpod_requirements.txt`](../../tools/gpu_backtest_tools/runpod/requirements.txt).

## Results

Expand Down
10 changes: 5 additions & 5 deletions docs/runpod.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ remote job, and state file. Server error bodies/auth headers are not printed.
From an installed checkout:

```bash
gpu-backtest runpod --config examples/rsi.json \
gpu-backtest runpod --config examples/gpu_backtest_examples/rsi/config.json \
--ssh-key ~/.ssh/runpod_ed25519 --output-dir runs/runpod-rsi --charts
```

Expand All @@ -48,7 +48,7 @@ The output directory must be new or empty. A dry run reads no API credentials,
requires no SSH key, and rents no GPU:

```bash
gpu-backtest runpod --config examples/rsi.json \
gpu-backtest runpod --config examples/gpu_backtest_examples/rsi/config.json \
--output-dir runs/runpod-preview --dry-run
```

Expand Down Expand Up @@ -97,7 +97,7 @@ runpod/pytorch:2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04
This existing official image supplies Python 3.11, SSH, and CUDA development
libraries. The engine does not use its PyTorch installation. The launcher creates
an isolated venv and installs the engine with the pinned
[validated environment](../src/gpu_backtest/runpod_requirements.txt), using the
[validated environment](../tools/gpu_backtest_tools/runpod/requirements.txt), using the
image's CUDA toolkit. Optional charts use Altair 6.3.0. Overrides must supply Python 3.11–3.13,
compatible CUDA development libraries/driver, SSH, tar, and `nvidia-smi`.

Expand Down Expand Up @@ -151,7 +151,7 @@ source .venv/bin/activate
python -m pip install -e '.[dev,viz,cuda]' 'cuda-bindings>=12.9.1,<13'
NUMBA_ENABLE_CUDASIM=0 gpu-backtest gpu-check --output runs/gpu_check.json
NUMBA_ENABLE_CUDASIM=0 gpu-backtest pipeline \
--config examples/rsi.json --output-dir runs/example
--config examples/gpu_backtest_examples/rsi/config.json --output-dir runs/example
gpu-backtest charts --input runs/example/*_top_entry.csv runs/example/*_top_exit.csv \
--output runs/example/charts.html
```
Expand All @@ -161,5 +161,5 @@ launcher deletes only pods it creates. Small hardware checks establish GPU
compilation and explicit numeric cases, not large-grid speed or every strategy's
correctness.

See [the hardware validation record](runpod-validation.md) for the tested image,
See [the hardware validation record](../benchmarks/results/runpod_validation_20261003.md) for the tested image,
driver, package versions, numeric results, and successful cleanup.
Loading
Loading