openapi: 3.1.0
info:
  title: Komplex AI Hallucination Detector
  version: 0.1.1
  summary: LLM hallucination detection over HTTP.
  description: |
    The Komplex AI API scores LLM responses for hallucination risk and
    returns a calibrated probability plus a breakdown across hallucination
    regimes (fabrication, fake citation, self-contradiction, etc.).

    All endpoints are JSON over HTTPS. The headline endpoint
    `POST /api/detect` accepts an LLM response (and an optional prompt)
    and returns the score. The remaining endpoints manage API keys and
    surface current-period usage.

    **Server-side use only.** API keys (`sk_*`) must never be embedded
    in browser-side JavaScript — they grant access to your account's
    quota and billing. Always call this API from a backend that holds
    the key in an environment variable. See `docs/api.md` ("API
    Security") for the recommended proxy pattern.

    **Zero retention.** The detector does not persist prompts or
    responses. Only structured telemetry (request_id, latency, units
    billed, regime, status code) is recorded for billing and operational
    use. See the Privacy Policy for full details.
  contact:
    name: Komplex AI Support
    email: support@komplexai.io
    url: https://komplexai.io
  license:
    name: Commercial
    url: https://komplexai.io/terms

servers:
  - url: https://api.komplexai.io
    description: Production
  - url: http://localhost:3000
    description: Local development (Next.js)

security:
  - bearerAuth: []

tags:
  - name: Detect
    description: Hallucination detection.
  - name: Keys
    description: API key management. Requires a signed-in Komplex AI dashboard session — these endpoints are not callable with a Bearer API key.
  - name: Usage
    description: Account usage and billing summary. Requires a signed-in Komplex AI dashboard session.

paths:
  /api/detect:
    post:
      tags: [Detect]
      operationId: detect
      summary: Score an LLM response for hallucination risk.
      description: |
        Submit an LLM-generated response (optionally with the prompt that
        produced it) and receive a calibrated hallucination probability
        plus per-regime scores.

        **Modes.**
        - `response`-only (omit `prompt`): higher precision on
          self-contained factual claims; safe default for the public web
          UI.
        - `prompt`+`response`: higher accuracy when context matters
          (e.g. "the model answered question X with Y"). Billed by total
          input size — see units note below.

        **Units / billing.** Each call is billed in 2,048-character
        input windows (`ceil(input_chars / 2048)`), where `input_chars`
        is the sum of `prompt` and `response` lengths. The billed amount
        is returned as `detections_billed` on the response.

        **Idempotency.** Pass the same `X-Request-Id` on a retry to
        collapse duplicate writes to the audit log (the second call will
        not produce an extra `usage_events` row).

        **Latency.** Warm-path p95 target is 5s; the proxy enforces a
        30s hard timeout and returns `502` if the upstream detector is
        unreachable.

        **Cold start (free tier).** The detector scales to zero when idle
        to keep the free tier free, so the *first* call after a quiet
        period warms the model up and can take ~10-30s; later calls are
        fast (sub-second). If a first call times out or returns `503`,
        wait a few seconds and retry. This is current free-tier behavior
        and improves as usage grows.
      parameters:
        - in: header
          name: X-Request-Id
          required: false
          description: |
            Optional client-supplied request ID. Used for idempotent
            retries and log correlation across hops. If absent, the
            server mints one and returns it on the response header of
            the same name. Format is free-form ASCII (recommended: an
            opaque token, e.g. a ULID).
          schema:
            type: string
            maxLength: 128
            example: req_a1b2c3d4e5f6a7b8
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/DetectRequest'
            examples:
              responseOnly:
                summary: Response-only mode (no prompt)
                value:
                  response: |
                    The Eiffel Tower was built in 1889 by Gustav Eiffel
                    to commemorate the centennial of the French
                    Revolution.
              promptPlusResponse:
                summary: Prompt + response mode (higher accuracy)
                value:
                  prompt: When was the Eiffel Tower built and why?
                  response: |
                    The Eiffel Tower was built in 1889 by Gustav Eiffel
                    to commemorate the centennial of the French
                    Revolution.
                  top_k_regimes: 6
      responses:
        '200':
          description: Detection succeeded.
          headers:
            X-Request-Id:
              description: |
                Server-minted request ID for log correlation. Always
                present on every response, success or error.
              schema:
                type: string
                example: req_a1b2c3d4e5f6a7b8
            X-Client-Request-Id:
              description: |
                Echoed back from the request's `X-Request-Id` header if
                the client supplied one. Lets the caller stitch its own
                client-side logs to our server logs.
              schema:
                type: string
            X-Units-Billed:
              description: |
                Detection units consumed by this call (`= detections_billed`
                in the response body). `0` on non-200 responses. At v1
                a single short call is one unit.
              schema:
                type: integer
                minimum: 0
                example: 1
            X-Quota-Period:
              description: |
                Current billing period anchor as `YYYY-MM`. Phase A:
                calendar month. Stripe (M3) will switch to the
                subscription period anchor.
              schema:
                type: string
                example: '2026-05'
            X-Quota-Remaining:
              description: |
                Best-effort units remaining in the current billing period.
                Omitted if the lookup fails (the audit row is already
                written; the header is a courtesy). Free-plan default is
                10,000 units/month.
              schema:
                type: integer
                minimum: 0
                example: 8166
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DetectResponse'
              examples:
                noHallucination:
                  summary: Low-probability — likely truthful
                  value:
                    p_hallucination: 0.08
                    flag: false
                    top_regime: NORMAL
                    regime_scores:
                      - { regime: NORMAL,        p: 0.72 }
                      - { regime: FABRICATED,    p: 0.11 }
                      - { regime: Other,         p: 0.07 }
                      - { regime: NEAR_FALSE,    p: 0.07 }
                      - { regime: CF_AUTH,       p: 0.02 }
                      - { regime: FALSE_REFUSAL, p: 0.01 }
                    request_id: req_a1b2c3d4e5f6a7b8
                    detections_billed: 1
                    mode: short
                    latency_ms: 412
                    model_version: nl-v1
                    calibrator_version: platt-v3
                    input_mode_used: ro
                    task_used: multiclass
                hallucinationFlagged:
                  summary: Hallucination flagged — fabricated claim
                  value:
                    p_hallucination: 0.87
                    flag: true
                    top_regime: FABRICATED
                    regime_scores:
                      - { regime: FABRICATED,    p: 0.61 }
                      - { regime: CF_AUTH,       p: 0.18 }
                      - { regime: NEAR_FALSE,    p: 0.09 }
                      - { regime: NORMAL,        p: 0.06 }
                      - { regime: Other,         p: 0.05 }
                      - { regime: FALSE_REFUSAL, p: 0.01 }
                    request_id: req_b2c3d4e5f6a7b8c9
                    detections_billed: 1
                    mode: short
                    latency_ms: 388
                    model_version: nl-v1
                    calibrator_version: platt-v3
                    input_mode_used: pr
                    task_used: multiclass
        '400':
          description: |
            Malformed request: missing `response`, body is not valid
            JSON, input exceeds the 2,048-character cap per field, or a
            Pydantic / pipeline validation error from the upstream
            detector (e.g. `top_k_regimes` / `task` / `input_mode` out
            of range). The response body conforms to `ErrorResponse`
            when emitted by the Vercel proxy; upstream FastAPI
            validation pass-through uses the `{"detail": "..."}` shape.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                missingResponse:
                  summary: '`response` field missing or empty (proxy)'
                  value:
                    error: response_required
                badJson:
                  summary: Request body is not valid JSON (proxy)
                  value:
                    error: bad_json
                inputTooLong:
                  summary: '`response` or `prompt` exceeds 2,048-char v1 cap (proxy)'
                  value:
                    error: input_too_long
                    field: response
                    max_chars: 2048
                    provided_chars: 3104
                    message: >-
                      v1 supports response inputs up to 2048 characters.
                      Long-document detection is planned for a later release.
                upstreamDetail:
                  summary: Upstream FastAPI validation error (pass-through)
                  value:
                    detail: "input_mode must be 'pr' or 'ro', got 'xyz'"
        '401':
          description: |
            Authentication required or invalid. Returned when this route
            is configured to require an API key and the
            `Authorization: Bearer sk_*` header is missing, malformed,
            or references a revoked key. The body conforms to
            `ErrorResponse`. The response sets
            `WWW-Authenticate: Bearer realm="komplexai"`.
          headers:
            WWW-Authenticate:
              schema:
                type: string
                example: Bearer realm="komplexai"
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error: unauthorized
                message: >-
                  This endpoint requires an API key. Send your sk_*
                  token as `Authorization: Bearer <key>`. Create a key
                  at /account/keys.
                request_id: req_c3d4e5f6a7b8c9d0
        '402':
          description: |
            Quota exceeded with no overage allowed. Applies to all hard-cap
            plans: `api_free` (3,700/mo), `web_free` (300/mo), `web_pro`
            (6K/mo + 200/day), `web_premium` (12K/mo + 300/day). Paid API
            tiers (`starter` / `growth` / `scale`) allow overage and never
            return 402 — those are billed via the daily Stripe overage cron.
            The `quota_type` field discriminates monthly vs daily; daily
            caps reset at 00:00 UTC. Response includes `upgrade_url`
            pointing at the billing page.
          headers:
            X-Plan:
              schema: { type: string }
              description: Internal plan name that produced the 402.
            X-Quota-Type:
              schema: { type: string, enum: [monthly, daily] }
              description: Whether the breached cap is monthly or daily.
            X-Quota-Limit:
              schema: { type: integer }
              description: The breached cap (units).
            X-Quota-Period-Used:
              schema: { type: integer }
              description: Units consumed in the breached period.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                monthlyExceeded:
                  summary: Monthly quota hard-stop (api_free / web_*)
                  value:
                    error: quota_exceeded
                    quota_type: monthly
                    plan: api_free
                    message: Free-tier monthly quota reached. Upgrade for more capacity.
                    units_used: 3700
                    quota_units: 3700
                    upgrade_url: https://detector.komplexai.io/pricing
                dailyExceeded:
                  summary: Daily quota hard-stop (web_pro 200/day or web_premium 300/day)
                  value:
                    error: quota_exceeded
                    quota_type: daily
                    plan: web_pro
                    message: Daily quota of 200 detections reached for the web_pro plan. Resets at 00:00 UTC. Upgrade for higher daily caps or wait until tomorrow.
                    units_used: 200
                    quota_units: 200
                    upgrade_url: https://detector.komplexai.io/pricing
        '403':
          description: |
            Account suspended. The user has been suspended for violating
            the Komplex AI terms of use (typically via the abuse-defense
            email-classifier pipeline, or by an admin). All `/api/detect`
            calls return 403 until the suspension is lifted on appeal. The
            response includes `appeal_url`.
          headers:
            X-Account-Status:
              schema:
                type: string
                example: suspended
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error: account_suspended
                message: >-
                  This account has been suspended for violating the
                  Komplex AI terms of use. To appeal, visit the appeal
                  form referenced below.
                reason: threat_high
                appeal_url: https://detector.komplexai.io/account/appeal
                request_id: req_e5f6a7b8c9d0e1f2
        '429':
          description: |
            Rate limited. The caller is sending requests faster than the
            tier permits. Inspect `Retry-After` (seconds) and back off.
          headers:
            Retry-After:
              schema:
                type: integer
                example: 1
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error: rate_limited
                message: Too many requests; retry after 1 second.
        '500':
          description: |
            Server-side failure. Two flavors:

            - `usage_audit_unavailable` — the audit-log writer
              (Postgres) was unable to record this request and the
              server refused to ship an un-audited response. The
              request was *not* served and is *not* billed.
            - `handler_failed` — the route handler threw unexpectedly.
              The audit row is still written (`status_code=500`).

            Retry with a fresh `X-Request-Id`.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                auditUnavailable:
                  summary: usage_events insert failed
                  value:
                    error: usage_audit_unavailable
                    message: >-
                      Could not record usage for this request. The
                      request was not served to protect billing
                      integrity. Please retry.
                    request_id: req_d4e5f6a7b8c9d0e1
                handlerFailed:
                  summary: Handler threw unexpectedly
                  value:
                    error: handler_failed
                    request_id: req_e5f6a7b8c9d0e1f2
        '502':
          description: |
            Upstream detector unreachable. The Next.js proxy could not
            contact the inference backend (network, 30s timeout, or
            non-HTTP failure). Retry after a brief delay.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error: upstream_unreachable
        '503':
          description: |
            Detector is not configured (the proxy is missing
            `ULM_API_URL` or `INFERENCE_SECRET`) or the upstream model
            is still warming up. Callers should retry shortly.
            `GET /health` (on the inference backend) will indicate
            readiness.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              examples:
                notConfigured:
                  summary: Proxy env not configured (proxy)
                  value:
                    error: detector_not_configured
                modelUnavailable:
                  summary: Upstream model not yet loaded (pass-through)
                  value:
                    detail: model_unavailable

  /api/keys:
    post:
      tags: [Keys]
      operationId: createApiKey
      summary: Create a new API key.
      description: |
        Mint a new API key for the signed-in dashboard user. The
        plaintext key is returned **exactly once** in the response body
        — there is no retrieval endpoint, only the hash is stored.

        This endpoint requires a Komplex AI **dashboard session**
        (the cookie set by `https://komplexai.io/account`). It is *not*
        callable with a Bearer API key.
      security:
        - sessionCookie: []
      requestBody:
        required: false
        content:
          application/json:
            schema:
              type: object
              additionalProperties: false
              properties:
                name:
                  type: string
                  maxLength: 64
                  description: |
                    Optional human-readable label, shown in the
                    dashboard. Whitespace is collapsed and the value is
                    trimmed to 64 characters.
                  example: production-server
      responses:
        '201':
          description: Key created. The plaintext `key` is returned once.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ApiKeyCreateResponse'
              example:
                id: 9c5f6a7b-8c9d-0e1f-2a3b-4c5d6e7f8a9b
                name: production-server
                key: sk_A1B2c3D4e5F6g7H8i9J0k1L2m3N4o5P6q7R8s9T0u1V2w3X4y5Z6
                prefix: sk_A1B2
                created_at: '2026-05-12T18:30:00Z'
        '401':
          description: Not signed in.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error: unauthorized
                message: unauthorized
    get:
      tags: [Keys]
      operationId: listApiKeys
      summary: List API keys for the current user.
      description: |
        Return all of the signed-in user's API keys (metadata only —
        never the plaintext). By default, revoked keys are omitted; pass
        `include_revoked=true` to include them.
      security:
        - sessionCookie: []
      parameters:
        - in: query
          name: include_revoked
          required: false
          schema:
            type: boolean
            default: false
          description: |
            If `true`, include keys whose `revoked_at` is set. Default
            is `false`.
      responses:
        '200':
          description: Key list.
          content:
            application/json:
              schema:
                type: object
                required: [keys]
                properties:
                  keys:
                    type: array
                    items:
                      $ref: '#/components/schemas/ApiKey'
              example:
                keys:
                  - id: 9c5f6a7b-8c9d-0e1f-2a3b-4c5d6e7f8a9b
                    prefix: sk_A1B2
                    name: production-server
                    created_at: '2026-05-12T18:30:00Z'
                    last_used_at: '2026-05-12T19:02:11Z'
                    revoked_at: null
                  - id: 7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d
                    prefix: sk_Q9R8
                    name: staging
                    created_at: '2026-05-10T12:00:00Z'
                    last_used_at: null
                    revoked_at: null
        '401':
          description: Not signed in.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error: unauthorized
                message: unauthorized

  /api/keys/{id}:
    delete:
      tags: [Keys]
      operationId: revokeApiKey
      summary: Revoke an API key.
      description: |
        Soft-revoke a key by setting `revoked_at = NOW()`. The row stays
        in the database for audit; subsequent authentications with this
        key will fail. The endpoint is idempotent — a second call on an
        already-revoked key returns `404`, identical to the
        non-existent-key case (existence of other users' keys is never
        disclosed).
      security:
        - sessionCookie: []
      parameters:
        - in: path
          name: id
          required: true
          schema:
            type: string
            format: uuid
          description: The api_keys.id UUID.
      responses:
        '200':
          description: Key revoked.
          content:
            application/json:
              schema:
                type: object
                required: [revoked, id]
                properties:
                  revoked:
                    type: boolean
                    enum: [true]
                  id:
                    type: string
                    format: uuid
              example:
                revoked: true
                id: 9c5f6a7b-8c9d-0e1f-2a3b-4c5d6e7f8a9b
        '401':
          description: Not signed in.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error: unauthorized
                message: unauthorized
        '404':
          description: |
            Key not found, not owned by the signed-in user, or already
            revoked. All three cases return the same status to avoid
            disclosing the existence of other users' keys.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
              example:
                error: bad_request
                message: not_found

  /api/usage/summary:
    get:
      tags: [Usage]
      operationId: getUsageSummary
      summary: Current-period usage and 12-month history.
      description: |
        Return the signed-in user's plan, current billing period totals,
        and a 12-month history of monthly usage. Powers the `/account`
        dashboard.

        Billing periods are calendar months in UTC; the period anchor
        (`currentPeriod.start`) is the first of the current month.

        Phase A note: until the Clerk integration ships,
        `usage_events.user_id` is `NULL` for traffic recorded through
        the public proxy and the returned counts reflect *global*
        traffic, not strictly per-user. After Phase B the same endpoint
        filters by `user_id`.
      security:
        - sessionCookie: []
      responses:
        '200':
          description: Usage summary.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/UsageSummary'
              example:
                user:
                  email: dev@example.com
                  name: Devin Example
                  initials: DE
                  memberSince: '2026-04-01'
                plan:
                  name: Free
                  quota: 10000
                  overageRatePer1k: 0
                currentPeriod:
                  start: '2026-05-01'
                  requests: 1834
                  units: 1842
                  quota: 10000
                  pct: 18
                  periodEndLabel: Jun 1, 2026
                history:
                  - period: May 2026
                    periodStart: '2026-05-01'
                    requests: 1834
                    units: 1842
                    overage: 0
                    charge: $0.00
                  - period: Apr 2026
                    periodStart: '2026-04-01'
                    requests: 902
                    units: 905
                    overage: 0
                    charge: $0.00

components:
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: API Key (sk_*)
      description: |
        API key authentication. Pass your key as
        `Authorization: Bearer sk_<your_key>`.

        Keys are issued from the dashboard at `/account/keys` and look
        like `sk_<43 base64url characters>` (256-bit entropy). Treat
        them as secrets — never embed in browser-side JavaScript or
        commit them to version control. See `docs/api.md` for the
        recommended backend-proxy pattern.

        Revoked keys return 401; rotate via the dashboard.
    sessionCookie:
      type: apiKey
      in: cookie
      name: __session
      description: |
        Komplex AI dashboard session cookie (Clerk-issued in production;
        a development stub in local builds). Required for the
        `/api/keys/*` and `/api/usage/*` endpoints — these are
        dashboard-internal and are *not* part of the public Bearer-key
        surface.

  schemas:
    DetectRequest:
      type: object
      required: [response]
      additionalProperties: false
      properties:
        response:
          type: string
          minLength: 1
          maxLength: 2048
          description: |
            The LLM-generated text to score. Required. Capped at 2048
            characters at the free tier; longer inputs are rejected
            with `400 bad_request` (sliding-window long-doc support is
            paid-tier only and not yet exposed at launch).
        prompt:
          type: string
          maxLength: 2048
          description: |
            Optional. The prompt the LLM saw before producing
            `response`. Including the prompt enables prompt+response
            mode, which is more accurate when context matters. Both
            `prompt` and `response` lengths are summed for billing
            (`detections_billed = ceil((len(prompt) + len(response)) / 2048)`).
        top_k_regimes:
          type: integer
          minimum: 1
          maximum: 10
          default: 6
          description: |
            How many regime scores to return, sorted by descending
            probability. Default `6` — `NORMAL` plus the 5 publicly
            advertised regimes (`FABRICATED`, `NEAR_FALSE`, `CF_AUTH`,
            `FALSE_REFUSAL`) and `Other`. The model's raw multiclass
            head emits 7 raw classes; `SELF_CONTR` and `UNDERSPECIFIED`
            are collapsed into `Other` server-side per the v1
            public-class mapping. Canonical across the upstream FastAPI,
            the Vercel proxy at `/api/detect`, and the public API docs.
        task:
          type: string
          enum: [binary, multiclass]
          description: |
            Which classifier head to use. `binary` returns a single
            `binary_hallucination` regime; `multiclass` returns the
            full per-regime breakdown. Public `/api/detect` (Vercel
            proxy) defaults to `multiclass` when absent; the upstream
            FastAPI Pydantic default is `binary`.

    RegimeScore:
      type: object
      required: [regime, p]
      additionalProperties: false
      properties:
        regime:
          type: string
          description: |
            Regime code. Publicly advertised values for the v1 NL
            release: `NORMAL` (benign), `FABRICATED`, `NEAR_FALSE`,
            `CF_AUTH`, `FALSE_REFUSAL`, and `Other`. The raw model
            emits 7 classes; `SELF_CONTR` and `UNDERSPECIFIED` are
            collapsed into `Other` server-side per the v1
            public-class mapping. Binary-task responses use
            `binary_hallucination`. Treat this field as an
            open-string enum — new regimes may be added in future
            model releases.
          example: FABRICATED
        p:
          type: number
          format: double
          minimum: 0
          maximum: 1
          description: Calibrated probability for this regime. The full `regime_scores` array sums to approximately 1.0 (multiclass head).
          example: 0.61

    DetectResponse:
      type: object
      required:
        - p_hallucination
        - flag
        - top_regime
        - regime_scores
        - request_id
        - detections_billed
        - mode
        - latency_ms
        - model_version
        - calibrator_version
        - input_mode_used
        - task_used
        - warnings
      additionalProperties: false
      properties:
        p_hallucination:
          type: number
          format: double
          minimum: 0
          maximum: 1
          description: Calibrated probability that the response contains a hallucination.
          example: 0.87
        flag:
          type: boolean
          description: |
            Convenience flag — `true` when `p_hallucination` crosses
            the model-dependent decision threshold for the
            (`task_used`, `input_mode_used`) cell that served the
            call. The threshold is NOT a fixed `0.5`; it varies per
            cell (Youden-optimal on the v2 val set). The exact value
            is an implementation detail of the active model build —
            check `calibrator_version` / `/v1/health` for the
            deployed build. Clients that need a different operating
            point should re-threshold from `p_hallucination`
            directly.
          example: true
        top_regime:
          type: string
          description: |
            Code of the highest-probability regime in `regime_scores`.
            Note that `NORMAL` is a valid value here when the model
            is confident the response is benign.
          example: FABRICATED
        regime_scores:
          type: array
          minItems: 1
          maxItems: 10
          items:
            $ref: '#/components/schemas/RegimeScore'
          description: |
            All requested regime scores sorted by descending `p`. In
            multiclass mode, scores sum to approximately 1.0. In binary
            mode, the array contains a single entry with regime
            `binary_hallucination`.
        request_id:
          type: string
          description: |
            ID of this request, suitable for log correlation and
            idempotent retry. Echoes the `X-Request-Id` request header
            when supplied, otherwise minted server-side. Also returned
            in the `X-Request-Id` response header.
          example: req_a1b2c3d4e5f6a7b8
        detections_billed:
          type: integer
          minimum: 0
          description: |
            Billing units consumed by this request. One unit = a 2,048
            character input window
            (`ceil((len(prompt) + len(response)) / 2048)`).
            Always `0` on non-200 responses.
          example: 1
        mode:
          type: string
          enum: [short, long]
          description: |
            Inference mode. `short` is single-pass (input fits in one
            window). `long` is sliding-window over a longer input
            (paid-tier; not exposed at v1 launch). Reflects sliding
            behavior, NOT the classifier head — see `task_used` for
            the head that served the call.
          example: short
        latency_ms:
          type: integer
          minimum: 0
          description: Server-side inference latency in milliseconds (excludes network).
          example: 388
        model_version:
          type: string
          description: |
            Build identifier of the detector that served the request.
            Stable per Modal deployment.
          example: nl-v1
        calibrator_version:
          type: string
          description: |
            Identifier for the probability calibrator (e.g. Platt /
            isotonic build) applied on top of the raw model score.
          example: platt-v3
        input_mode_used:
          type: string
          enum: [pr, ro]
          description: |
            Which input cell actually served this call. `pr` =
            prompt+response (server saw both); `ro` = response-only
            (server never saw the prompt at all). Mirrors the
            `input_mode` request override when supplied, otherwise
            reports the server's auto-detect.
          example: pr
        task_used:
          type: string
          enum: [binary, multiclass]
          description: |
            Echoes the task that was actually used to produce
            `p_hallucination`. Equals the value of the request's
            `task` field if explicitly set; otherwise `binary` (the
            server default). The multiclass head is only used when
            `task="multiclass"` is supplied.
          example: multiclass
        warnings:
          type: array
          default: []
          items:
            type: string
            enum:
              - quota_80pct
              - quota_95pct
              - overage_active
              - calibration_uncalibrated
              - input_truncated
              - subscription_inactive
          description: |
            Advisory codes attached to this response. Always present;
            usually `[]`.

            Current codes (closed enum at the v1 launch surface):

            - `quota_80pct` — user has used ≥80% of the monthly
              quota (free + paid tiers both).
            - `quota_95pct` — ≥95%; last warning before hard-stop on
              free, or before overage on paid.
            - `overage_active` — paid tier with usage above the
              included limit; this call is billed at the overage
              rate.
            - `calibration_uncalibrated` — the active calibrator is
              an `uncalibrated:*` placeholder. Should only occur
              during deployment / model rollouts. Treat scores with
              elevated uncertainty.
            - `input_truncated` — input was truncated to fit the
              per-call character cap. (Reserved; not currently
              emitted by the v1 server.)
            - `subscription_inactive` — the user has a non-active
              subscription_status but is still calling the API.

            **Forward-compat:** clients MUST ignore unknown codes.
            Future model / billing revisions may add codes here
            without bumping the API version.
          example: []

    ApiKey:
      type: object
      required: [id, prefix, created_at]
      additionalProperties: false
      properties:
        id:
          type: string
          format: uuid
          description: api_keys.id (UUID). Use this when revoking via DELETE.
        prefix:
          type: string
          description: |
            First seven characters of the plaintext key
            (e.g. `sk_A1B2`). Safe to display; insufficient to
            authenticate.
          example: sk_A1B2
        name:
          type:
            - string
            - 'null'
          maxLength: 64
          description: Optional user-supplied label.
          example: production-server
        created_at:
          type: string
          format: date-time
          description: Creation timestamp (UTC).
          example: '2026-05-12T18:30:00Z'
        last_used_at:
          type:
            - string
            - 'null'
          format: date-time
          description: |
            Wall-clock UTC time the key was last successfully used to
            authenticate. `null` if never used.
          example: '2026-05-12T19:02:11Z'
        revoked_at:
          type:
            - string
            - 'null'
          format: date-time
          description: |
            UTC time the key was revoked, or `null` if active. Revoked
            keys are filtered out of `GET /api/keys` unless
            `include_revoked=true`.

    ApiKeyCreateResponse:
      type: object
      required: [id, key, prefix, created_at]
      additionalProperties: false
      properties:
        id:
          type: string
          format: uuid
        name:
          type:
            - string
            - 'null'
          maxLength: 64
        key:
          type: string
          description: |
            The plaintext API key. **Returned exactly once.** Store it
            securely (env var, secret manager). Format is
            `sk_<43 base64url characters>` with 256-bit entropy.
          example: sk_A1B2c3D4e5F6g7H8i9J0k1L2m3N4o5P6q7R8s9T0u1V2w3X4y5Z6
        prefix:
          type: string
          description: First seven characters of `key` (for later display).
          example: sk_A1B2
        created_at:
          type: string
          format: date-time
          example: '2026-05-12T18:30:00Z'

    UsageSummary:
      type: object
      required: [user, plan, currentPeriod, history]
      additionalProperties: false
      properties:
        user:
          type: object
          additionalProperties: false
          properties:
            email:
              type:
                - string
                - 'null'
              format: email
            name:
              type:
                - string
                - 'null'
            initials:
              type:
                - string
                - 'null'
            memberSince:
              type:
                - string
                - 'null'
              format: date
        plan:
          type: object
          required: [name, quota, overageRatePer1k]
          additionalProperties: false
          properties:
            name:
              type: string
              description: Plan name (e.g. `Free`, `Pro`). Free at launch.
              example: Free
            quota:
              type: integer
              minimum: 0
              description: Monthly unit quota included in the plan.
              example: 10000
            overageRatePer1k:
              type: number
              format: double
              minimum: 0
              description: |
                Per-1,000-unit overage rate in USD. `0` until Stripe
                wiring lands; the field is reserved so client code does
                not have to change shape when overage billing arrives.
              example: 0
        currentPeriod:
          type: object
          required: [start, requests, units, quota, pct, periodEndLabel]
          additionalProperties: false
          properties:
            start:
              type: string
              format: date
              description: First day of the current billing period (UTC).
              example: '2026-05-01'
            requests:
              type: integer
              minimum: 0
              description: Count of HTTP-200 detect calls so far this period.
              example: 1834
            units:
              type: integer
              minimum: 0
              description: Total billing units this period.
              example: 1842
            quota:
              type: integer
              minimum: 0
              description: Duplicate of `plan.quota` for client convenience.
              example: 10000
            pct:
              type: integer
              minimum: 0
              description: |
                Percent of quota consumed, rounded to nearest integer.
                Can exceed 100 once overage starts (Stripe phase).
              example: 18
            periodEndLabel:
              type: string
              description: |
                Human-friendly label for the first day of the next
                billing period — e.g. when the quota resets / when the
                next invoice is generated.
              example: Jun 1, 2026
        history:
          type: array
          maxItems: 12
          description: |
            Up to the last 12 months of usage, newest first. Each entry
            is the rollup of one billing period.
          items:
            type: object
            required: [period, periodStart, requests, units, overage, charge]
            additionalProperties: false
            properties:
              period:
                type: string
                description: Month label (`Mon YYYY`).
                example: May 2026
              periodStart:
                type: string
                format: date
                example: '2026-05-01'
              requests:
                type: integer
                minimum: 0
              units:
                type: integer
                minimum: 0
              overage:
                type: integer
                minimum: 0
                description: Overage units this period. `0` until Stripe wiring lands.
              charge:
                type: string
                description: |
                  Pre-formatted overage charge, e.g. `$0.00`. Returned
                  as a string so the UI does not need to know the
                  currency formatting rule.
                example: $0.00

    ErrorResponse:
      type: object
      additionalProperties: true
      description: |
        Common error envelope. Non-2xx responses emitted by the Vercel
        proxy and `withUsageTracking` wrapper conform to this shape.
        Upstream FastAPI errors that pass through unchanged use the
        `{"detail": "..."}` shape instead — clients should accept
        either (`error` || `detail`).

        Specific endpoints may add fields (e.g. `upgrade_url` for
        quota errors, `request_id` for any path that wrote to the
        audit log).
      properties:
        error:
          type: string
          enum:
            - bad_json
            - response_required
            - input_too_long
            - unauthorized
            - account_suspended
            - quota_exceeded
            - rate_limited
            - usage_audit_unavailable
            - handler_failed
            - upstream_unreachable
            - detector_not_configured
          description: |
            Machine-readable error class emitted by the Vercel proxy
            or `withUsageTracking` wrapper. Stable enum at v1; new
            codes are added under semver-compatible doc updates and
            never removed in-place. Pass-through upstream FastAPI
            errors use the `detail` field below instead.
        detail:
          type: string
          description: |
            FastAPI-style error string from the upstream detector
            (`detector_api/server.py`), passed through unchanged.
            Typical values: `invalid_inference_secret` (401),
            `model_unavailable` (503), `inference_failed` (503),
            `not_implemented: ...` (503), pipeline `ValueError`
            text (400).
        message:
          type: string
          description: Human-readable explanation. Safe to log; never contains caller content.
        quota_type:
          type: string
          enum: [monthly, daily]
          description: |
            Present when `error == quota_exceeded`. Disambiguates which
            cap was breached. Daily caps reset at 00:00 UTC.
        plan:
          type: string
          description: |
            Present when `error == quota_exceeded` or `account_suspended`.
            Internal plan slug that produced the response.
          example: web_pro
        units_used:
          type: integer
          description: |
            Present when `error == quota_exceeded`. Units consumed in
            the breached period.
        quota_units:
          type: integer
          description: |
            Present when `error == quota_exceeded`. The breached cap.
        field:
          type: string
          description: |
            Present when `error == input_too_long`. Which field exceeded
            the cap (`response` or `prompt`).
          enum: [response, prompt]
        max_chars:
          type: integer
          description: |
            Present when `error == input_too_long`. The character cap
            that was exceeded (2048 at v1).
        provided_chars:
          type: integer
          description: |
            Present when `error == input_too_long`. The actual character
            count that was provided.
        reason:
          type: string
          description: |
            Present when `error == account_suspended`. Internal reason
            code from the abuse-defense pipeline (e.g. `threat_high`,
            `threat_med_borderline`, `manual_admin`). Nullable.
        appeal_url:
          type: string
          format: uri
          description: |
            Present when `error == account_suspended`. Points to the
            user-facing appeal form.
          example: https://detector.komplexai.io/account/appeal
        upgrade_url:
          type: string
          format: uri
          description: |
            Present when `error == quota_exceeded`. Points to the
            billing page where the caller can raise their plan limit.
          example: https://komplexai.io/account/billing
        request_id:
          type: string
          description: |
            Present on errors that were nonetheless audited
            (e.g. 401/500 from the usage-tracking wrapper). Use for
            support / log correlation.
          example: req_d4e5f6a7b8c9d0e1
