# Phase 0 cloud GPU POC runbook

Status: Phase 0 technical POC passed with production blockers

Last verified: 2026-08-14

## Scope

This is the single-instance Phase 0 proof of concept from the 90-day AI
execution plan. It validates one self-hosted text model before any Kubernetes,
multi-node, RAG, training, or customer-facing rollout.

## Inventory

| Item | Verified value |
|---|---|
| Region / zone | Kuala Lumpur `ap-southeast-3a` |
| ECS instance | `i-8psbjd6ru30b79uxd2xb` / `peta-ai-kl-l20-poc` |
| Instance type | `ecs.gn8is.2xlarge`, 8 vCPU, 64 GiB RAM |
| GPU | 1 x NVIDIA L20, 46,068 MiB visible to the OS |
| OS | Ubuntu 22.04, kernel `5.15.0-177-generic` |
| GPU driver | `580.126.09` |
| Docker | `29.5.1` |
| NVIDIA Container Toolkit | `1.17.8-1` |
| System disk | 100 GiB ESSD PL0; 8.0 GiB free after image, model, and compile cache |
| Billing | Pay-as-you-go; reference compute price USD 2.641736/hour |
| Public network | Traffic billing, 5 Mbps peak; model API port is not open |
| Security group | Public ICMP and TCP/22; no HTTP, HTTPS, or vLLM ingress |

The console is authoritative for the current public IP. It can change after a
stop/start cycle because the instance does not have an EIP.

## Access

SSH uses the dedicated local key:

```bash
ssh -i ~/.ssh/peta-ai-kl-l20-poc_ed25519 petaops@CURRENT_PUBLIC_IP
```

Verified SSH policy:

- `petaops` has passwordless sudo.
- Public-key authentication is enabled.
- Password and keyboard-interactive authentication are disabled.
- Root SSH login is disabled.
- Alibaba Cloud Assistant remains the recovery path if SSH becomes unavailable.

The security group still allows TCP/22 from `0.0.0.0/0`. Restrict it to an
approved stable office/VPN CIDR after confirming that the source address will
not change unexpectedly.

## Model service

Deployment files live in `/opt/peta-ai` on the server and
`deploy/ai-poc` in the repository.

Pinned artifacts:

- vLLM `v0.23.0`, amd64 image digest
  `sha256:3a1e7f5904e1a1192a02aa0086ceaffc33985d7044c7bb25b3a43d61bdbe3ac0`.
- `Qwen/Qwen3.6-27B-FP8`, revision
  `e89b16ebf1988b3d6befa7de50abc2d76f26eb09`.
- Text-only serving, 32K context, four concurrent sequences. The original 16K
  benchmark conditions are retained below as historical evidence.

The service listens only on `127.0.0.1:8000`. Do not add a security-group rule
for port 8000.

The host secret is stored as `VLLM_SERVICE_KEY` in `/opt/peta-ai/.env`.
Compose maps it to the container's `VLLM_API_KEY` environment variable. Do not
pass it with the vLLM `--api-key` option or share unredacted container-inspection
output; both can expose the value.

### Start and inspect

```bash
cd /opt/peta-ai
sudo docker compose config --quiet
sudo docker compose up -d
sudo docker compose ps
sudo docker compose logs --tail=200 vllm
curl --fail http://127.0.0.1:8000/health
```

### Stop and restart the model

```bash
cd /opt/peta-ai
sudo docker compose stop vllm
sudo docker compose up -d
```

Stopping the container does not stop ECS billing.

### Call the API on the server

Load the secret without printing it:

```bash
cd /opt/peta-ai
set -a
. ./.env
set +a
python3 benchmark.py --concurrency 1 --requests 3
```

This command measures non-thinking mode, which matches the latency-sensitive
lead-triage JSON use case. Add `--enable-thinking` only for a separately labeled
reasoning benchmark.

## Verified results

Test conditions: 2026-08-10, one L20, vLLM `v0.23.0`, official Qwen3.6-27B-FP8
revision shown above, 16K context, four execution slots, and short English
real-estate lead JSON prompts. Reported throughput counts completion tokens.

| Workload | Requests | TTFT p50 / p95 | Latency p50 / p95 | Aggregate output tokens/s |
|---|---:|---:|---:|---:|
| Non-thinking, concurrency 1 | 5 | 0.136 / 0.136 s | 4.450 / 4.450 s | 17.34 |
| Non-thinking, concurrency 2 | 8 | 0.226 / 0.605 s | 4.213 / 4.640 s | 35.65 |
| Non-thinking, concurrency 4 | 12 | 0.192 / 0.337 s | 4.240 / 4.429 s | 71.51 |
| Non-thinking, concurrency 8 with four slots | 16 | 4.326 / 4.367 s | 8.374 / 8.422 s | 73.10 |
| Non-thinking stability loop, concurrency 4 | 40 | 0.192 / 0.193 s | 4.247 / 4.265 s | 72.47 |
| Thinking mode, concurrency 1 | 1 | 0.089 / 0.089 s | 14.577 / 14.577 s | 17.56 |

The five-request concurrency-1 run peaked at 38,751 MiB VRAM, 100% GPU
utilization, and 246 W. The concurrency-4 run also peaked at 38,751 MiB and
100% utilization, with 262 W peak power. The 40-request stability loop
completed in 42.447 seconds with no error, exception, failure, or OOM keywords
in the corresponding server logs.

Functional checks passed for authenticated chat completion, streaming usage,
valid constrained JSON, and `qwen3_coder` tool-call parsing. `/health` returns
200 without authentication; `/v1/models` returns 401 without the service key
and 200 with it. The model port was verified as bound to `127.0.0.1` and an
external connection to public port 8000 timed out.

Initial model download took 2,468 seconds for a 28.75 GiB checkpoint. Weight
loading took 18.25 seconds and used 27.64 GiB before KV-cache allocation. The
benchmark server reported 9.65 GiB of GPU KV cache, 137,216 token capacity, and
an 8.38x theoretical cache concurrency at 16,384 tokens. A cached container
restart returned `/health` 200 after 101 seconds; it loaded weights in 5.03
seconds, reused compilation artifacts in 6.79 seconds, and expanded the final
KV cache to 11.71 GiB / 166,570 tokens (10.17x at 16K). A four-request,
four-way post-restart smoke test and the subsequent health check passed.

After an ECS Economical Mode stop/start on 2026-08-10, the guest booted at
20:46:06, the container started at 20:46:15, and vLLM reported application
startup complete at 20:51:00. This is a 4 minute 54 second guest-boot recovery
(4 minutes 45 seconds from container start). The cold weight load took 224.83
seconds for 27.64 GiB; cached compilation took 11.53 seconds. The resulting KV
cache was 11.79 GiB / 167,936 tokens (10.25x at 16K). Health, authenticated
model listing, and a non-thinking chat smoke test all passed with no container
restart or OOM; the chat returned `PETA_OK` in 0.258 seconds. The public IP did
not change in this cycle, but operators must still recheck it after every
Economical Mode start.

vLLM reported that several L20 FP8 matrix shapes used the default W8A8 kernel
configuration because tuned files were unavailable. The results above are a
valid unoptimized baseline, not a hardware throughput ceiling.

### Long-transcript recovery validation (2026-08-14)

The service was moved from 16,384 to 32,768 maximum model tokens without
changing the model revision, analysis prompt, or JSON schema. vLLM reported a
145,635-token GPU KV cache and 4.44x theoretical concurrency at 32,768 tokens,
which remains above the configured four execution slots. A cached restart
reached authenticated model-list health in about 150 seconds with zero
container restarts.

Three previously failed PropertyLab analyses were then rerun serially through
the real petaV3 job path. The transport used one attempt with a 240-second
request timeout, preventing a still-running GPU request from being duplicated
after the old 60-second timeout.

| Meeting ID | Input / output tokens | Provider time | First-pass JSON / schema | Estimated GPU cost |
|---:|---:|---:|---:|---:|
| 351 | 19,399 / 982 | 67.155 s | Pass / Pass | USD 0.0493 |
| 333 | 11,686 / 1,490 | 90.787 s | Pass / Pass | USD 0.0666 |
| 329 | 10,893 / 1,217 | 74.621 s | Pass / Pass | USD 0.0548 |

All three completed in one attempt, with no OOM or container restart. The
PropertyLab projection ended at 13 done, zero failed, and zero processing rows.
The combined USD 0.1707 figure is an allocation estimate using provider time
and USD 2.641736/hour; it is not an Alibaba Cloud invoice. Two results still
exceeded the 60-second product target, so this fixes reliability but does not
claim the P95 latency gate has passed.

### Anonymized business-shaped evaluation

With user approval, test shapes were derived from `PETA - Requirements.docx`
and recent property-consultation recording patterns in Google Drive. No raw
transcript, name, contact detail, or customer identifier was sent to the GPU.
Four concurrent, synthetic agency requests covered a lead sales brief, call
analysis, Property Advisor response, and WhatsApp follow-up. Unique isolation
markers were supplied to each request and instructed not to be repeated.

The four-request run completed 1,256 output tokens in 21.521 seconds (58.36
aggregate output tokens/s). All requests returned HTTP 200 and no response
contained its own marker or another agency's marker. The sales brief and
WhatsApp draft passed their structural checks. The call analysis hit the
initial 384-token ceiling and produced incomplete JSON; it passed on rerun with
a 768-token ceiling, finishing at 498 tokens in 26.977 seconds. Output limits
therefore need to be workload-specific rather than globally fixed at 384.

The first Property Advisor response conflated title-form and tenure concepts.
A rerun with an explicit Malaysian property glossary correctly separated
freehold/leasehold tenure, residential/commercial category, and
strata/individual title form, but still introduced general financing and cost
claims not present in the supplied context. This is a **No-Go for autonomous
customer-facing advice** until retrieval, source attribution, claim-level
grounding checks, and human review are implemented. The isolation-marker check
is only an inference-request smoke test; it does not validate tenant
authorization, persistence, logging, cache keys, or application data isolation.

### Local SSH tunnel

For an operator test from an approved workstation:

```bash
ssh -N -L 18000:127.0.0.1:8000 \
  -i ~/.ssh/peta-ai-kl-l20-poc_ed25519 \
  petaops@CURRENT_PUBLIC_IP
```

This tunnel does not expose port 8000 publicly.

### petaV3 application integration

petaV3 has a first-class, company-managed `peta` AI provider that reuses the
existing OpenAI-compatible `AiClient` transport. It is separate from the real
OpenAI provider, so changing the self-hosted endpoint cannot redirect GPT
traffic. The provider defaults to `peta-qwen3.6-27b-fp8`, sends
`chat_template_kwargs.enable_thinking=false`, requests usage in the final SSE
chunk, and supports the existing encrypted key storage, request logging,
prompt/model pins, retries, and circuit breaker. It is text-only and global-only
in Phase 0; members cannot view or enter the service key.

For local integration, start the tunnel first and set non-secret environment
values to match its local port:

```dotenv
PETA_AI_BASE_URL=http://127.0.0.1:18000
PETA_AI_MODEL=peta-qwen3.6-27b-fp8
PETA_AI_REQUEST_TIMEOUT=240
PETA_AI_RETRY_ATTEMPTS=1
```

Then clear the Laravel config cache. In **Manage → Integrations → AI Requests →
AI Providers**, an administrator must personally paste the vLLM service key into
the **PETA AI (Self-hosted)** card. Saving verifies `GET /v1/models`; the
subsequent **Test** action sends a tiny chat request and records provider, model,
tokens, duration, request purpose, and success/failure in `ai_requests` without
logging the key. Only after both checks pass should an administrator pin an
individual prompt to `peta`, or set `AI_DEFAULT_PROVIDER=peta` for the whole
environment.

### Live petaV3 business-workload evidence (2026-08-10)

The `ai_conversation` prompt was pinned to `peta-qwen3.6-27b-fp8` and exercised
through the real user-portal streaming UI. The request used one existing,
non-personal Project Catalogue row (`Iris Residence @ Mont Kiara`) and accurately
stated that its stored median price, median PSF, and property type were missing.
No customer record was copied because the local test account had no Wealth Plan
or Property Analysis, and the local Zoom dataset contained zero recordings.

| Metric | Result |
|---|---:|
| Status | Success |
| Input / output tokens | 1,736 / 1,127 |
| End-to-end duration | 64.7 s |
| Average output rate | 17.4 tokens/s |
| First visible content | Observed between 1 and 6 s; exact TTFT not instrumented |
| Marginal API cost logged | USD 0.0000; ECS cost remains hourly |

The reply correctly identified missing data, did not invent a specific price,
rent, return, or buy/sell verdict, and followed the PropertyLab identity and
2+2+1 call-to-action. It was nevertheless too long for an interactive answer,
rendered its Markdown table poorly, and made some location- and property-type
inferences more confidently than the supplied record justified. This is a
technical integration pass but only a **conditional quality pass**. Before a
customer pilot, shorten the prompt/output budget and add a stricter rule that
catalogue facts not present in the supplied context must be labeled as general
knowledge or unknown. `zoom_recording_chat` is pinned to PETA for the next
manual test, but remains unverified until an approved recording exists locally.

Do not use an SSH tunnel as the production topology. A deployed petaV3 instance
must reach the model through a private network or authenticated TLS gateway;
vLLM port 8000 remains closed to the public Internet. Phase 0 also retains the
existing free-credit policy: a member can use the company key while credits
remain, but production subscription/billing behaviour for a global-only
self-hosted provider is not yet defined.

## Logs and storage

```bash
cd /opt/peta-ai
sudo docker compose logs --since=30m vllm
sudo docker stats --no-stream peta-ai-vllm
nvidia-smi
df -h /
du -sh models-cache
sudo docker system df
```

The Hugging Face cache is `/opt/peta-ai/models-cache`; the compilation cache is
`/opt/peta-ai/vllm-cache`. Do not delete either during a routine container
restart. The system disk is intentionally small, so only one 27B quantization
and the active vLLM image should be retained.

## Recovery

1. Check ECS state, current public IP, disk usage, and the security group in the
   Alibaba Cloud console.
2. Use Cloud Assistant if SSH is unavailable.
3. Validate `sshd` before reloading it: `sudo sshd -t`.
4. Check `nvidia-smi`, then repeat the CUDA container smoke test:

   ```bash
   sudo docker run --rm --gpus all \
     nvidia/cuda:12.8.1-base-ubuntu22.04 nvidia-smi
   ```

5. Check `sudo docker compose logs --tail=300 vllm` for model download, OOM,
   quantization, or context-length failures.
6. If the FP8 model cannot meet the memory/concurrency gate, test a reviewed,
   checksum-pinned AWQ build; do not silently replace the official model.

## Cost guardrails

At the reference compute price, continuous runtime is approximately USD 63.40
per day. The USD 500 Month 1 gate is reached after roughly 189 compute hours
(7.9 days) before disk and network charges.

At the measured four-way stability throughput, compute alone is approximately
USD 10.13 per million output tokens. This excludes prompt processing, idle
time, storage, network, operations, and application overhead, so it is not a
customer price.

- Budget alerts at 50%, 80%, and 100% are not yet verified.
- OOS scheduled stop with `StoppedMode=StopCharging` is not yet configured or
  verified.
- Automatic release is disabled.
- Release protection is currently disabled.
- Stopping, releasing, resizing, or purchasing resources always requires human
  approval. An OS-level `shutdown` is not accepted as evidence of StopCharging.

## Phase 0 evidence

| Gate | Result |
|---|---|
| Instance configuration matches order | Pass |
| SSH key and sudo access | Pass |
| Root/password SSH disabled | Pass |
| Host and container GPU visibility | Pass |
| Pinned vLLM image pull | Pass |
| Qwen3.6-27B-FP8 startup | Pass |
| OpenAI-compatible API, JSON, stream, and tool call | Pass |
| TTFT / latency / tokens per second | Pass for the recorded short-prompt workload |
| Concurrency and short stability | Pass through four active sequences; queueing visible at eight requests |
| Anonymized business-shaped four-agency workload | Pass after workload-specific output-limit tuning |
| Cross-agency inference marker smoke test | Pass; application and persistence isolation remain unverified |
| Autonomous customer-facing property advice | No-Go pending grounded retrieval, citations, and human review |
| petaV3 `peta` provider adapter | Pass in feature tests for verify, chat, stream, usage, and hidden member key |
| Live petaV3 → tunneled L20 request | Pass: real portal stream, 1,736 / 1,127 tokens in 64.7 s |
| Long PropertyLab transcripts | Reliability pass at 32K: three prior failures recovered in one attempt; latency gate still pending |
| 9B comparison baseline | Pending |
| Budget alerts and StopCharging drill | Pending human approval |

Decision: **conditional Go for continued Phase 0 experimentation on this one
L20; No-Go for Phase 1 or customer production deployment.** The technical gate
passes for short JSON and tool-calling workloads, and four-way batching gives
near-linear aggregate throughput. Before any broader rollout, resolve the
8.0 GiB disk margin, restrict SSH to a stable approved CIDR, run the 9B cost
baseline plus long-context grounded evaluations, implement application-level
tenant-isolation tests, tune or quantify the default FP8 kernels, and complete
budget alerts plus a proven StopCharging drill.

## Addendum — PropertyLab evaluation metric formulas (2026-08-11 branch)

The in-app Evaluation dashboard (`/manage/zoom/evaluation`, super admin only)
computes its figures with these exact formulas; anything it cannot verify is
shown as unavailable, never zero:

| Figure | Formula | Verified from |
|---|---|---|
| Human review coverage | saved section reviews ÷ (completed PropertyLab generations × 4 sections) | review + generation tables |
| Score averages / quality pass | arithmetic means; pass = all three scores ≥ 4 | saved reviews only |
| Human-confirmed unsupported rate | `Issue found` ÷ (`Issue found` + `No issue`); `Not sure` + unanswered excluded, shown separately | tri-state review answers |
| Automatic flag rate (diagnostic) | flagged (contradicted/no-evidence) ÷ evaluated transcript-verifiable units; judgment/external counted separately | evaluator verdicts (server-validated evidence) |
| First-pass JSON / schema rate | strict raw-response checks ÷ queue attempts with a provider response | attempt reliability flags (runtime only — backfilled rows are `unknown`) |
| P95 end-to-end | nearest-rank over `completed_at − queued_at`; unstable < 20 samples | generation timings |
| Failure rate | settled failed ÷ all settled generations | generation statuses |
| Cloud evaluator cost | Σ evaluator `ai_requests` estimated cost | logged token usage × configured provider/model price table (**Estimate**) |
| PropertyLab GPU cost | hourly price × request seconds ÷ 3600 — **Estimate**, may double-count concurrency | configured price × timings (estimate) |
| Benchmark allocated cost | verified billable window ÷ completed generations | **only** a human-verified benchmark run (unverified otherwise) |

Evidence fields that remain **unverified** on this branch: GPU price inputs
(no billing integration — operator-entered), benchmark billable windows
(manual designation), and any figure sourced from backfilled lineage rows
(first-pass history is recorded as unknown). GPU infrastructure and ECS
lifecycle changes stay manual, human-confirmed operations.
