Benchgen

Env Diagnostics Benchmark

1 phaseActive

Diagnostic benchmark that echoes EVERY environment variable the run container receives (secrets masked) into the logs — for both the ingestion and scoring stages. Use it to confirm that env injection works, that the DIAG_ prefix is applied with no key conflicts, and (optionally) that the injected model credentials actually work. Env-var-first: no model.py required.

Overview

Env Diagnostics — Overview

This is a diagnostic benchmark. It does not measure model quality. Instead it echoes every environment variable the run container receives, so you can confirm that the platform's env-var injection works end to end.

What it checks

  1. Declared DIAG_* variables arrived — the model connection (auto-filled from the selected model) and any runner-typed values.
  2. No key conflicts — every variable is listed with its group, so duplicate or clashing keys are obvious.
  3. Both stages receive the env — the ingestion and scoring containers each print their own DIAG_* set.
  4. (Optional) The model credentials work — set DIAG_PING_MODEL=true to make one live chat request and see the reply in the logs.

Where to look

  • Logs — grouped list of every variable (secrets masked) for ingestion and scoring. This is the primary output.
  • Detailed Results — an HTML table of every variable with its group and whether it was present.
  • Score Breakdown:
    • env_health — % of expected declared keys that arrived.
    • diag_vars / total_vars — counts.
    • model_ping — 1 if the live model call succeeded.
    • both_stages — 1 if DIAG_* vars reached both containers.

Variables

KeySourceNotes
DIAG_API_URLmodelauto-filled model endpoint
DIAG_API_KEYmodel (secret)auto-filled bearer key
DIAG_MODEL_NAMEmodelauto-filled model name
DIAG_TEMPERATURE / DIAG_MAX_TOKENS / DIAG_TIMEOUTmodelsampling
DIAG_NOTErunnerfree text echoed in the logs
DIAG_PING_MODELrunnertrue to attempt one live model call

Secrets are always masked in both the logs and the HTML report.