THETA AI SERVICES
On-demand Model APIs
Want to deploy your dedicated model? Click here
API Offerings
- API (12)
- WebUI (7)
- LLM (4)
- ImageGen (3)
- MoE (2)
- ImageCaption (1)
- ImageRestore (1)
- ObjectDetection (1)
- SpeechRec (1)
- VideoGen (1)
Model Details
gpt-oss-120b
- Type: LLM
- API: Yes
- MoE: Yes
The gpt‑oss‑120b model is a 117‑billion‑parameter, open‑weight mixture‑of‑experts language model by OpenAI, released under the permissive Apache 2.0 license. Its architecture features 128 experts per layer, with only 4 activated per token, delivering precise and efficient performance. It supports chain‑of‑thought reasoning with configurable effort levels, native tool use (web browsing, function calling, code execution), structured outputs, and an impressive 128 K‑token context window. Benchmarks show it rivals or outperforms OpenAI’s proprietary o4‑mini on tasks in reasoning, coding, health, and expert domains.
- Pricing: $0.04 / 1M input tokens, $0.20 / 1M output tokens
- View
Blip
- Type: WebUI, API, ImageCaption
BLIP is a vision-language model from Salesforce that excels at image captioning. It integrates a Vision Transformer and a BERT-like encoder-decoder, jointly optimizing contrastive learning, image-text matching, and language modeling. This enables it to produce accurate, context-rich captions and achieve state-of-the-art results on image captioning benchmarks.
- Pricing: $0.01 / image
- View
Qwen3 Parallax
- Type: LLM, API
Alibaba’s Qwen3-32B-FP8 and Parallax come together on Theta EdgeCloud to enable a new paradigm for decentralized AI inference. We have adapted Parallax’s distributed serving architecture to run seamlessly across Theta’s global community node network, enabling large models like Qwen3 32B to be hosted with pipeline parallelism over the internet—without reliance on centralized GPU clusters. This approach harnesses heterogeneous edge compute to deliver scalable, cost-efficient, and low-latency inference, unlocking a resilient and geographically distributed foundation for real-world AI applications.
- Pricing: $0.20 / 1M input tokens, $0.40 / 1M output tokens
- View
Whisper
- Type: SpeechRec, API
The OpenAI Whisper model is a robust, versatile automatic speech recognition (ASR) system designed to transcribe and understand audio from diverse sources and languages with high accuracy. It is built to recognize speech in noisy environments, distinguish between different speakers, and even translate languages, making it a powerful tool for unlocking the potential of audio data across various applications.
- Pricing: $2.50 / 1M tokens
- View
Stable Diffusion XL Turbo
- Type: ImageGen, WebUI, API
The Stable Diffusion model is a cutting-edge AI tool designed for generating highly detailed and creative images from textual descriptions, enabling artists and creators to bring their visions to life with unprecedented ease and flexibility. It utilizes advanced deep learning techniques to understand and interpret user prompts, producing visual content that ranges from realistic to fantastical, thereby revolutionizing the field of digital art and content creation.
- Pricing: $0.01 / image
- View
LLaVA
- Type: LLM, ImageCaption, ObjectDetection
LLaVA represents a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding, achieving impressive chat capabilities mimicking spirits of the multimodal GPT-4 and setting a new state-of-the-art accuracy on Science QA.
- Pricing: $0.01 / image
- View