Enterprise AI Infrastructure Bangkok, Thailand

The global token factory

TAPEX is an AI cloud built for one job: generating tokens at scale. Latest-generation NVIDIA compute, high-throughput inference, and cost-per-token economics — served from Thailand to the world.

Verified customers · Aligned with US export controls (EAR/BIS)

Blackwell
Latest-generation NVIDIA compute
Multi-tier
GPU fleet, workstation to rack scale
In-region
Latency & data residency for ASEAN
24/7
Operations & support
The platform

Built on the frontier of accelerated computing

TAPEX deploys NVIDIA's latest-generation accelerated computing as its foundational compute layer — engineered for inference performance per watt and per dollar.

Foundational compute layer

NVIDIA GB300 Grace Blackwell at the core

Our clusters pair Grace CPUs with Blackwell-generation GPUs on a coherent memory fabric — purpose-matched to large-model inference and fine-tuning, and surrounded by a multi-tier GPU fleet so every workload runs on right-sized silicon.

Flagship platformNVIDIA GB300 Grace Blackwell
Fleet tiersRTX PRO 6000 · GeForce RTX 5090
Optimised forHigh-volume inference · Fine-tuning
LocationThailand · ASEAN region

Blackwell-generation performance

Early access to GB300-class hardware puts TAPEX at the frontier of tokens per second — with RTX PRO 6000 and GeForce RTX 5090 tiers for workloads that don't need rack scale.

Serving-optimised stack

Continuous batching, paged KV-cache and per-model quantisation tuned across the fleet — throughput without quality loss.

Performance per watt

Efficient power and cooling design translates directly into lower, more predictable cost per token.

Token-factory economics

A factory measured in tokens, not racks

We run our infrastructure the way a factory runs a production line — utilisation, yield and unit cost are the metrics that matter.

Unit economics

Optimised cost per token

Every layer of the stack — hardware selection, batching, scheduling, power — is tuned toward one number: what a million tokens costs to produce.

High utilisation

Dense scheduling and workload mixing keep GPUs producing instead of idling — utilisation is our core discipline.

Efficient power & cooling

Facility design tuned for Blackwell-class density, so energy goes into tokens — not overhead.

Throughput, guaranteed

Capacity reservations with clear throughput commitments — plan token volumes like any other supply contract.

Transparent supply

A trusted, verified counterparty in the NVIDIA channel — customers know exactly where their compute comes from.

Services

Compute, the way you need it

Inference-as-a-service

High-throughput serving for production LLM workloads, priced per token.

  • OpenAI-compatible endpoints
  • Batch & real-time
  • Per-token pricing

GPU-as-a-service

Dedicated Blackwell-class GPU capacity, bare-metal control when you need it.

  • Reserved instances
  • Full-node allocation
  • Custom images

Fine-tuning

Adapt open and custom models on large-memory systems built for long context and large batches.

  • LoRA & full fine-tunes
  • Private datasets stay in-region

Dedicated capacity

Long-term reserved clusters for model providers and enterprises with sustained token demand.

  • Committed throughput
  • Predictable unit cost
The production line

How a token factory works

A factory turns raw inputs into finished goods. Ours turns energy, silicon and model weights into tokens — the unit every AI product is built from.

Raw inputs

Power, cooling and Blackwell-class silicon — provisioned, dense and ready. The plant floor of the factory.

Models loaded

Open-weight and customer models resident in GPU memory, quantised and compiled for the hardware they run on.

Orchestration

Continuous batching and smart scheduling mix workloads to keep every GPU producing — utilisation is yield.

Tokens delivered

Streamed through low-latency APIs, metered per token — with throughput you can plan against like any supply contract.

Energy+Silicon+Models=Tokens at scale

Models we serve

The open-weight frontier, on tap

Production-grade serving for the model families teams actually deploy — or bring your own weights and we'll run them.

Llama

Meta

The default open-weight family for chat, agents and enterprise assistants.

ChatAgents

DeepSeek

DeepSeek

Frontier-class reasoning at open-weight economics — built for hard problems.

ReasoningCode

Qwen

Alibaba

Strong multilingual coverage — including Thai and Southeast Asian languages.

MultilingualVision

Mistral

Mistral AI

Efficient dense and MoE models with excellent cost-to-quality ratios.

ChatEfficient

Gemma

Google

Lightweight models that shine on high-volume, latency-sensitive workloads.

LightweightFast

GPT-OSS

OpenAI

Open-weight reasoning models with broad ecosystem and tooling support.

ReasoningAgents

Bring your own model

Custom checkpoints, fine-tunes and proprietary architectures — deployed on dedicated capacity, weights and data kept in-region.

Deploy your model
Why ASEAN

Your inference layer, in your region

Hyperscalers serve the world from somewhere else. TAPEX serves it from Thailand — with a home-ground advantage for Southeast Asia: a regional alternative to Western neoclouds, built for the workloads and regulations of this market.

Lower latency for ASEAN users

In-region serving cuts round-trips to distant clouds — critical for real-time products.

Data residency

Workloads and datasets stay in Thailand, simplifying regional data-governance requirements.

Built for regional scale

Capacity for AI startups, enterprises and model providers across Thailand and ASEAN.

Trust & governance

A counterparty you can verify

Advanced compute carries responsibility. TAPEX operates with full transparency toward NVIDIA channel partners, regulators and customers.

Customer verification

Every customer completes KYC before access. We know who runs on our infrastructure — always.

End-use verification

Workloads are screened against declared end uses, with ongoing monitoring for diversion risk.

US export-control alignment

Full alignment with EAR/BIS regulations — a trusted, transparent counterparty for NVIDIA channel partners and enterprise customers.

Start producing tokens at scale

Tell us about your workload — model, token volumes, latency needs — and our team will come back with capacity and pricing.