# Odds acquisition worksheet

An offline Python 3.10+ planning model for comparing a batched API with an authorized collector on the **same required workload**. It does not contact bookmakers, test an API, verify a permission, assess an odds feed, or select a winner. No external packages are needed.

Put these five downloaded files in one directory: `odds_acquisition.py`, `example.json`, `blank.json`, `test_odds_acquisition.py`, and `README.md`. Open a terminal in that directory and run:

```sh
python3 odds_acquisition.py example.json
python3 odds_acquisition.py blank.json
python3 -m unittest -v test_odds_acquisition.py
```

The first two commands print deterministic JSON. The blank worksheet prints unknown quantities and `hold`, with exit code 0 because it is valid input. Invalid values return exit code 2. Tests run entirely locally. To retain a result:

```sh
python3 odds_acquisition.py example.json > result.json
```

Copy `blank.json` to another filename, edit the JSON in a text editor, then pass that filename to the script. Keep `example.json` unchanged if running the included example tests.

## Record requirements before costs

`workload.description` should state the exact events, bookmakers, markets, settlement periods, use/storage requirements, tolerated source age and historical date range/resolution. The same scope must apply to both candidates. `units_per_poll` describes how each source serves that scope: two API batch requests can cover the same work as twenty page requests. It is **not** a bookmaker-count multiplier applied to both indiscriminately.

For each candidate, complete the four `gates`: `coverage`, `allowed_use`, `freshness` and `history`. Each has a status (`yes`, `no` or `unknown`) and a required evidence note. Write the actual requirement, dated document or authorized observation, and its limits in that note. If history is unnecessary, say so explicitly. A gate marked `no` or `unknown` produces `hold`. A note's existence is checked; its truth is not. Evidence is supplied by the reader, not inferred by the program.

`eligible_for_source_trial` means only that the entered gates, quantities, costs and aggregate quota permit further evaluation. It is not verified coverage, freshness, permission, cost completeness, or a recommendation to collect. Gates are shared across scenarios: if a scenario changes the actual requirement, copy the worksheet and reassess its gates. In particular, a shorter poll interval does not establish fresher source data.

## Inputs and units

- `active_days`: whole active days in the comparison/quote period. All costs and quotas must cover that same period.
- `hours_per_day`: a fixed daily window between 0 and 24 hours.
- Each scenario supplies `poll_interval_seconds` (positive) and `collector_billed_bytes_per_attempt` (whole bytes, including everything the relevant provider bills). The latter can differ from observed wire bytes.
- `units_per_poll`: whole HTTP requests per poll covering the complete required workload. Include browser subresources or other required calls in this count when appropriate; the example does not use a browser.
- `planned_retry_fraction`: extra attempts as a fraction of scheduled requests; 0.05 means add 5%, not an observed 5% failure probability. Retries use the same credit/byte assumptions as other attempts. Vary inputs if their billing differs; this simple model cannot represent separate success/failure tariffs.
- API `credits_per_attempt`: billable units per attempted request under your endpoint/market/region and retry assumptions. Credits are not dollars. Confirm actual billing; this fixed-unit model does not reproduce every provider's exceptions.
- Collector `wire_bytes_per_attempt`: optional known whole wire bytes, separately reported. Cost uses **billed** bytes from the scenario, not this field. The larger-billing sensitivity holds wire bytes constant to show that billing basis must be explicit.
- `quota`: `basis` is `requests` or (API only) `credits`; `limit` is the aggregate quote-period allowance. `unlimited: true` and `limit: null` explicitly assert no aggregate cap. Otherwise a null limit means unknown. Per-second limits/concurrency are not modeled.
- `costs`: `fixed_cost` (subscription/infrastructure/storage and other known fixed costs); `usage_price_per_unit` (currency per API credit or per collector **decimal GB**); `maintenance_hours`; and `hourly_rate`. Use one currency and comparison period throughout.

Use JSON `null` for unknown numeric inputs. An explicit `0` means known zero. Unknown multiplied by zero remains unknown; the worksheet never silently fills a missing input with zero. Numeric strings, booleans, negatives, nonfinite numbers, individual numbers above 10^18 and fractional request/byte counts are rejected. This magnitude bound excludes implausibly huge planning inputs. Non-object JSON, an invalid schema version and an unreadable input file also return exit code 2 without results. Notes must be nonempty. The input is a small documented schema, not a general quote parser: preserve required field names and replace their values.

## Arithmetic and interpretation

The sampling schedule polls at the start of each daily window, then every interval, excluding the window's ending instant:

```text
polls per day = ceil(hours_per_day × 3600 / poll_interval_seconds)
polls = active_days × polls per day
scheduled requests = polls × units_per_poll
planned retries = ceil(scheduled requests × planned_retry_fraction)
total attempts = scheduled requests + planned retries
API credits = total attempts × credits_per_attempt
collector billed bytes = total attempts × collector_billed_bytes_per_attempt
collector decimal GB = billed bytes / 1,000,000,000
collector GiB = billed bytes / 1,073,741,824
cost = fixed cost + usage units × usage price + maintenance hours × hourly rate
```

Ceiling preserves partial final intervals and whole retry attempts. This is a fixed sampling plan, not an event-aware scheduler or a forecast of usable records. It does not model latency, bandwidth capacity, overlapping work, partial payloads, authentication, parsing, changing event schedules, rate limits, historical backfill requests or a retry queue. Add separate modeled work where those are required.

The fictional API quote includes its credit allowance, so its incremental per-credit input is explicitly zero **within the quota**. Do not add the subscription cost twice. Tiered pricing or paid overage needs a fresh valid quote/input. When quota is exceeded or unknown, `modeled_total` is null even if the components are numerically known; the supplied quote is not established for this workload. Missing cost components also make the total null. A failed coverage gate can still display known arithmetic, but its decision remains `hold`.

`modeled_total` includes only supplied fixed, usage and recurring maintenance costs. The fictional collector's fixed cost includes infrastructure/storage. The example excludes one-time implementation/migration, taxes and any extra egress/incident costs not covered by those fictional inputs. Add period-allocated amounts to `fixed_cost`, adjust maintenance and state the allocation in `costs.note` for your comparison. This is neither a total-ownership guarantee nor a profitability model.

## Fictional worked example

Every price, coverage/permission assessment and freshness assumption in `example.json` is invented. No source was trialled. Both candidates serve the same 20-page event workload for 30 active days, eight hours a day, under an assumed 120-second source-age budget and no historical-backfill requirement. API batches use two requests per poll and three credits per attempt. The collector uses twenty requests per poll. Both budget 5% extra attempts. The supplied monthly API quote caps credits at 100,000; the collector assumes no aggregate request cap.

| Scenario | API attempts / credits | API modeled total | Collector attempts / billed GB | Collector modeled total |
| --- | ---: | ---: | ---: | ---: |
| 60 seconds, 200,000 billed bytes/attempt | 30,240 / 90,720 | $200 | 302,400 / 60.48 | $560.96 |
| 10 seconds, 200,000 billed bytes/attempt | 181,440 / 544,320 | Unknown: quota exceeded | 1,814,400 / 362.88 | $1,165.76 |
| 60 seconds, 1,000,000 billed bytes/attempt | 30,240 / 90,720 | $200 | 302,400 / 302.4 | $1,044.80 |

The baseline API total is $100 subscription + $0 incremental usage + 2 hours × $50. The collector is $40 fixed + 60.48 GB × $2 + 8 hours × $50. Lower polling intervals multiply requests even when no new odds become available. Bigger billed responses increase collector usage without changing its request count. These sensitivities demonstrate why different request batching and billing units must be recorded before comparing quoted totals; they do not establish that one real source is cheaper.
