# petaV3 Cloud GPU Internal AI Engineering SOP

Version: v1.1

Date: 2026-08-06

Audience: petaV3 engineers, DevOps, AI infrastructure owner
Goal: Deploy a self-hosted language model on cloud GPU and connect it to petaV3 as a testable, observable, rollback-safe `internal_ai` provider.

## 1. Purpose

This SOP is an engineering execution document. It answers:

1. How to prepare the cloud GPU server.
2. How to deploy vLLM and an open-weight model.
3. How to protect the model API with an AI Gateway.
4. How to add an `internal_ai` provider to petaV3.
5. How to migrate the first business feature.
6. How to prepare for RAG.
7. What to check before rollout.
8. How to debug and rollback.

The first implementation covers text AI only. Local video generation, image generation, TTS, and local transcription models are out of scope for the first engineering phase.

## 2. Target Architecture

```text
petaV3 Laravel / Vue
  |
  | AiClient -> internal_ai provider
  v
AI Gateway / Nginx
  |
  | OpenAI-compatible API
  v
vLLM Model Server
  |
  +-- Qwen / other open-weight model
  +-- NVIDIA GPU runtime
  +-- Docker
  |
  v
Qdrant Vector DB  (Phase 1 RAG)
```

## 3. Execution Order

| Order | Workstream | Owner | Acceptance |
|---|---|---|---|
| 1 | Provision cloud GPU server | Infra | `nvidia-smi` works |
| 2 | Install Docker + NVIDIA runtime | Infra | Docker container can access GPU |
| 3 | Deploy vLLM + model | Infra / AI Engineer | `/v1/chat/completions` works |
| 4 | Deploy AI Gateway | Infra | API key and IP allowlist work |
| 5 | Add `internal_ai` provider to petaV3 | Backend | Provider test succeeds |
| 6 | Migrate first business feature | Backend / Business | Real samples can run |
| 7 | Add logs, evaluation, cost tracking | Backend / Infra | POC report can be produced |
| 8 | Add RAG in Phase 1 | Backend / Data | Private data retrieval works |

## 4. Cloud GPU Server Preparation

### 4.1 Recommended Configuration

| Item | POC Recommendation | Notes |
|---|---|---|
| Cloud and region | Alibaba Cloud International, Kuala Lumpur `ap-southeast-3` first | Singapore `ap-southeast-1` is inventory / capability fallback only |
| Preferred GPU | `ecs.gn8is.2xlarge`, L20 48GB | Official specification: 8 vCPU, 64 GiB, one L20; check live stock before order |
| Fallback GPU | A10 24GB-class instance | Run a validated Qwen3-14B AWQ build; do not assume regional availability |
| Disk | 300-500GB encrypted ESSD | Model cache and logs; size Qdrant separately in Phase 1 |
| OS | Ubuntu 22.04 LTS | Fix one OS version for the first POC |
| Network | VPC; port 443 only from petaV3 | SSH via bastion / VPN / fixed IP; never expose 8000 |

Kuala Lumpur belongs to the Global Data Plane, not a Mainland China data plane. CN-test may use only synthetic or approved de-identified data until a separate Mainland region, account, and compliance review exist.

### 4.2 Account, Budget, and Order Gates

```text
[ ] Company Alibaba Cloud International account; no personal account
[ ] Least-privilege RAM operator with MFA
[ ] USD 500 POC cloud-spend hard cap
[ ] 50% / 80% / 100% budget alerts configured and tested
[ ] Live calculator / order quote and inventory evidence saved
[ ] Disk, snapshots, and EIP may continue billing after GPU stop
[ ] Primary=Kuala Lumpur and fallback=Singapore approval recorded
```

### 4.3 Server Initialization Checklist

```text
[ ] SSH key configured
[ ] Root login policy confirmed
[ ] SSH source IP restricted
[ ] Server timezone confirmed
[ ] Disk size sufficient
[ ] GPU visible
[ ] Docker running
[ ] Docker can access GPU
```

### 4.4 GPU Validation

```bash
nvidia-smi
```

Expected:

```text
GPU model, VRAM, driver version, and CUDA version are visible.
```

Docker GPU validation:

```bash
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
```

Expected:

```text
The same GPU is visible inside the container.
```

### 4.5 Economical Stop Outside Working Hours

Use an OOS schedule or ECS API `StopInstance` with `StoppedMode=StopCharging`. Running Linux `shutdown`, `poweroff`, or `halt` does not enter Alibaba Cloud economical mode.

Acceptance:

```text
[ ] OOS or controlled API owns working-hour start / stop
[ ] Console shows economical mode after stop, not only an OS shutdown
[ ] Billing distinguishes released GPU/CPU/RAM from continuing disk/EIP charges
[ ] Runbook records that restart may fail when GPU inventory is unavailable
[ ] Singapore or alternate-GPU recovery steps exist
```

## 5. vLLM Model Server Deployment

### 5.1 Directory Layout

Recommended:

```text
/opt/peta-ai/
  docker-compose.yml
  .env
  logs/
  models-cache/
  qdrant/
```

### 5.2 First Model Choice

| Model | Best For | Recommendation |
|---|---|---|
| Qwen 7B | Fast smoke test | Low-cost test |
| Qwen3-14B BF16 | First L20 business POC | Recommended for structured extraction |
| Qwen3-14B AWQ | A10 24GB fallback | Must be quality-compared with BF16 |
| Qwen 32B quantized | Later higher-quality analysis | Test only if 14B misses the target |

Start with Qwen3-14B and validate structured budget, objection, preference, intent, and next-action fields. Benchmark bilingual WhatsApp tone separately against the current frontier provider; draft quality alone must not invalidate a successful extraction POC.

### 5.3 vLLM Docker Example

```bash
docker run -d \
  --name peta-ai-vllm \
  --runtime nvidia \
  --gpus all \
  -v /opt/peta-ai/models-cache:/root/.cache/huggingface \
  -p 127.0.0.1:8000:8000 \
  --ipc=host \
  --restart unless-stopped \
  vllm/vllm-openai:<PINNED_TAG>@sha256:<APPROVED_DIGEST> \
  --model Qwen/Qwen3-14B \
  --revision <PINNED_MODEL_REVISION> \
  --served-model-name peta-qwen3-14b \
  --api-key "<VLLM_SERVICE_KEY>" \
  --host 0.0.0.0 \
  --port 8000
```

Notes:

- `-p 127.0.0.1:8000:8000` binds the model API to localhost only.
- Public or cross-server access should go through the AI Gateway.
- `<PINNED_TAG>`, digest, model revision, and key come from the deployment manifest / secret manager. Do not execute placeholders unchanged or commit secrets.
- Formal POC deployments prohibit floating `latest`; record model revision, tokenizer revision, and quantization checksum.

### 5.4 API Test

Run on the GPU server:

```bash
curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer <VLLM_SERVICE_KEY>" \
  -d '{
    "model": "peta-qwen3-14b",
    "messages": [
      {"role": "user", "content": "Summarize what a real estate sales copilot should do."}
    ]
  }'
```

Acceptance:

```text
[ ] API returns a response
[ ] Response contains an assistant message
[ ] Simple prompt returns in 3-10 seconds
[ ] vLLM logs show no OOM
```

## 6. AI Gateway / Nginx

### 6.1 Why Gateway Is Required

petaV3 should not call raw vLLM directly. The Gateway handles:

1. API key validation.
2. IP allowlist.
3. HTTPS.
4. Rate limits.
5. Timeouts.
6. Access logs.
7. Future model routing.

### 6.2 Minimal Gateway Flow

```text
client -> https://ai.company.com/v1/chat/completions
  -> Nginx checks API key and IP
  -> proxy_pass http://127.0.0.1:8000
```

### 6.3 Gateway Acceptance

```text
[ ] Missing API key returns 401/403
[ ] Wrong API key returns 401/403
[ ] Non-allowlisted IP cannot access
[ ] petaV3 server can access
[ ] Access log records requests
[ ] vLLM port 8000 is not publicly exposed
[ ] Publicly trusted certificate or private mTLS is used
[ ] petaV3 keeps TLS verify=true
[ ] No self-signed certificate plus verify=false bypass
```

## 7. petaV3 `internal_ai` Provider Integration

### 7.1 Existing AI Architecture

Current petaV3 AI call chain:

```text
Business feature
  -> AiClient
  -> AiCredential / provider config
  -> Transport
  -> Provider API
  -> AiResponse
  -> ai_requests log
```

Files to understand:

```text
src/Ai/Services/AiClient.php
src/Ai/AiCredential.php
src/Ai/Transports/*
src/Ai/Responses/AiResponse.php
config/ai_prompts.php
```

### 7.2 `internal_ai` Design

`internal_ai` should behave like an OpenAI-compatible provider:

```text
provider: internal_ai
base_url: https://ai.company.com/v1
api_key: company internal key
model: Qwen/Qwen3-14B
transport: OpenAI-compatible
```

### 7.3 Expected Code Areas

Exact files depend on the current branch, but expected areas include:

| Area | Work |
|---|---|
| AiCredential provider catalog | Add `internal_ai` |
| AI provider settings UI | Allow key / model configuration |
| Transport | Reuse OpenAI transport or add InternalAiTransport |
| Model list | Add internal model choices |
| Provider test | Support `internal_ai` test request |
| Config | Add base URL and default model |
| Tests | Provider validation, test request, AiClient response |

### 7.4 Provider Test Acceptance

```text
[ ] Admin can save internal_ai
[ ] Provider test returns OK
[ ] ai_requests records provider=internal_ai
[ ] Bad key produces clear error
[ ] Timeout produces clear error
[ ] Feature can switch back to Gemini / OpenAI
```

## 8. First Business Pilot

### 8.1 Recommended Priority

| Priority | Feature | Reason |
|---|---|---|
| 1 | ConversationAnalyzer | Structured extraction best validates the 14B model and real estate labels |
| 2 | AI Conversations | Additional chat / stream validation, not the only business evidence |
| 3 | WhatsApp Draft | High value, but must start with human review |

### 8.2 ConversationAnalyzer Pilot Steps

1. Freeze 20-50 blind-evaluation transcripts from a larger approved pool; do not replace hard examples after prompt tuning.
2. Label budget, objections, preferences, intent, and next action, then link known appointment, viewing, booking, SPA, and won/lost outcomes. Store unknown outcomes as `unknown`.
3. Run the current provider and save versioned outputs.
4. Run `internal_ai` and save versioned outputs.
5. Compare JSON stability and field-level precision / recall / F1.
6. Use at least two independent business reviewers and adjudicate disagreements.
7. Record latency, errors, human corrections, and cost per task.

Acceptance:

```text
[ ] Complete schema returned
[ ] Budget, intent, preferences, pain points extracted
[ ] Next action is actionable
[ ] JSON/schema success >= 95%
[ ] Human accuracy on critical fields >= 85% and no worse than baseline
[ ] Known outcome labels >= 90% complete in the blind evaluation set
[ ] Severe hallucination / data leakage = 0
```

### 8.3 WhatsApp Draft Guardrail

The first version must be draft mode only.

Reasons:

- Avoid wrong promises.
- Avoid inappropriate tone.
- Reduce WhatsApp account risk.
- Allow measurement of human edit rate.

Record `shown / accepted / edited / sent / dismissed / executed`. An edit rate above 50% is not an automatic No-Go; risky commitments must be zero, and the first version must never auto-send.

## 9. Phase 1 RAG Integration

### 9.1 First Data Sources

```text
call_recordings.transcript
customer_journey_reports
whatsapp_messages
leads
property_analyses
wealth_plans
video_projects.brief
video_projects.chat_state
```

### 9.2 Qdrant Collection Suggestion

First version can use one collection:

```text
collection: peta_v3_documents
vector_size: depends on embedding model
distance: cosine
```

Payload:

```json
{
  "source_type": "call_transcript",
  "source_id": 123,
  "lead_id": 456,
  "admin_id": 789,
  "user_id": null,
  "company_id": "propertylab",
  "created_at": "2026-07-08",
  "permission_scope": "sales_team",
  "title": "Call transcript with Lead #456"
}
```

### 9.3 RAG Permission Principle

Do not rely on prompts to prevent data leakage. Filtering must happen server-side.

Flow:

```text
User asks question
  -> petaV3 resolves allowed lead/team/company scope
  -> Qdrant search with metadata filter
  -> retrieved chunks filtered again in Laravel
  -> only allowed context sent to model
```

Acceptance:

```text
[ ] Sales user cannot retrieve another sales user's leads
[ ] Manager can retrieve team data
[ ] Admin can retrieve global data
[ ] External customer collections are fully isolated
```

## 10. Logging, Monitoring, and Cost

### 10.1 Required Logs

```text
provider
model
prompt key
subject type / id
lead id
request duration
status
error
input token estimate
output token estimate
retrieved document ids
```

### 10.2 GPU Monitoring

Minimum metrics:

```text
GPU utilization
VRAM usage
request count
latency p50 / p95
error rate
OOM count
container restart count
```

### 10.3 Cost Tracking

Daily tracking:

```text
GPU hourly cost
running hours
request count
average latency
estimated cost per request
continuing disk / snapshot / EIP cost
budget consumed percentage
```

Stop new experiments when POC cloud spend exceeds USD 500 and require project-owner review. Report engineering, labeling, and business-review labor separately from GPU pricing.

## 11. Security Rollout Checklist

```text
[ ] vLLM is not publicly exposed
[ ] Gateway requires API key
[ ] Gateway restricts source IP
[ ] HTTPS configured
[ ] TLS verify=true; no self-signed bypass
[ ] vLLM image tag + digest and model revision pinned
[ ] OOS / ECS API economical stop verified; Linux shutdown is not used as a substitute
[ ] USD 500 budget and alerts active
[ ] petaV3 secrets are not in git
[ ] ai_requests does not log API keys
[ ] RAG retrieval has permission filters
[ ] WhatsApp is draft mode only
[ ] Fallback follows regional policy; CN never automatically falls back to Global
[ ] Timeout configured
[ ] Errors are visible
[ ] Logs are available for debugging
```

## 12. Debug Checklist

### 12.1 vLLM Not Responding

Check:

```bash
docker ps
docker logs peta-ai-vllm --tail 200
nvidia-smi
curl http://127.0.0.1:8000/v1/models
```

Possible causes:

- Model still loading.
- GPU OOM.
- Container crashed.
- Port not bound.

### 12.2 petaV3 Cannot Call the Model

Check:

```text
base_url is correct
API key is correct
Gateway allowlists petaV3 server IP
Nginx access log
petaV3 laravel.log
ai_requests error
```

### 12.3 JSON Output Failure

Actions:

```text
Lower temperature
Strengthen schema prompt
Add retry
Add validator
Keep fallback provider
```

### 12.4 Poor Model Quality

Actions:

```text
Adjust prompt
Try larger model
Add RAG context
Add few-shot examples
Compare against Gemini output
Ask business users for concrete failure notes
```

## 13. Rollback Strategy

Every business feature using `internal_ai` must be rollback-safe:

```text
feature config: provider=internal_ai
rollback: provider=gemini/openai
```

Rollout principles:

1. Do not remove the original provider.
2. Switch each business feature independently.
3. Start with internal users.
4. Expand to a small group.
5. Roll out broadly only after evaluation.

## 14. Definition of Done

The first engineering phase is complete when:

```text
[ ] Cloud GPU model API runs stably
[ ] petaV3 internal_ai provider is configurable
[ ] Provider test succeeds
[ ] At least one business feature can call internal_ai
[ ] ai_requests logs are complete
[ ] latency / error / cost can be measured
[ ] fallback and rollback exist
[ ] POC evaluation data exists
```

After this, the project can proceed to Phase 1: RAG, more workflows, and internal case study.

## 15. Official References

- Alibaba Cloud region IDs: <https://www.alibabacloud.com/help/en/user-center/developer-reference/common-region-id-reference>
- Alibaba Cloud gn8is / L20 GPU specifications: <https://www.alibabacloud.com/help/en/ecs/user-guide/gpu-accelerated-compute-optimized-and-vgpu-accelerated-instance-families-1>
- Alibaba Cloud ECS economical mode: <https://www.alibabacloud.com/help/en/ecs/user-guide/economical-mode>
- vLLM OpenAI-compatible server: <https://docs.vllm.ai/en/latest/serving/online_serving/openai_compatible_server/>
- vLLM Docker deployment: <https://docs.vllm.ai/en/latest/deployment/docker/>
