{"object":"list","data":[{"id":"gemini-3.7-flash","object":"model","created":1786710605,"owned_by":"helixmind","display_name":"Gemini 3.7 Flash","description":"Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.","architecture":{"context_window":1048576,"max_output_tokens":384000,"modality":"text"},"pricing":{"prompt":0.00075,"completion":0.001875,"cache_read":0.0000375,"cache_write":0.000020830000000000002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"claude-sonnet-5-coding","object":"model","created":1785788823,"owned_by":"helixmind","display_name":"Claude Sonnet 5","description":"Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.","architecture":{"context_window":200000,"max_output_tokens":32000,"modality":"text"},"pricing":{"prompt":0.003,"completion":0.015,"cache_read":0.0003,"cache_write":0.00375},"supported_endpoints":["/v1/messages"]},{"id":"deepseek-v4-flash-0731-thinking","object":"model","created":1785692810,"owned_by":"helixmind","display_name":"DeepSeek V4 Flash 0731 (Thinking)","description":"DeepSeek V4 Flash 0731 Thinking enables DeepSeek's reasoning mode on the efficiency-optimized Mixture-of-Experts model with a 1M-token context window, built for fast inference, high-throughput workloads, reasoning, coding, and agent workflows.","architecture":{"context_window":1048576,"max_output_tokens":384000,"modality":"text"},"pricing":{"prompt":0.00014000000000000001,"completion":0.00028000000000000003,"cache_read":0.000028},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v4-flash-0731","object":"model","created":1785692747,"owned_by":"helixmind","display_name":"DeepSeek V4 Flash 0731","description":"DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.","architecture":{"context_window":1048576,"max_output_tokens":384000,"modality":"text"},"pricing":{"prompt":0.00014000000000000001,"completion":0.00028000000000000003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"],"default_params":{"thinking":{"type":"disable"}}},{"id":"claude-opus-4-6","object":"model","created":1785534766,"owned_by":"helixmind","display_name":"Claude Opus 4.6","description":"Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time. The model shows deeper contextual understanding, stronger problem decomposition, and greater reliability on hard engineering tasks than prior generations.","architecture":{"context_window":1000000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.005,"completion":0.025,"cache_read":0.0005,"cache_write":0.00625},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"claude-opus-5","object":"model","created":1785534494,"owned_by":"helixmind","display_name":"Claude Opus 5","description":"Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents.","architecture":{"context_window":1000000,"max_output_tokens":128000,"modality":"text"},"pricing":{"prompt":0.005,"completion":0.025,"cache_read":0.0005,"cache_write":0.00625},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemini-3.5-flash-lite","object":"model","created":1785450682,"owned_by":"helixmind","display_name":"Gemini 3.5 Flash Lite","description":"Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.","architecture":{"context_window":1000000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.0003,"completion":0.0025,"cache_read":0.000029999999999999997,"cache_write":0.00008333},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemini-3.1-pro-preview","object":"model","created":1785450581,"owned_by":"helixmind","display_name":"Gemini 3.1 Pro Preview","description":"Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation of the Gemini 3 series, it combines high-precision reasoning across text, image, video, audio, and code with a 1M-token context window. Reasoning Details must be preserved when using multi-turn tool calling, see our docs here: https://openrouter.ai/docs/use-cases/reasoning-tokens#preserving-reasoning","architecture":{"context_window":1000000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.002,"completion":0.012,"cache_read":0.0002,"cache_write":0.000375},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemini-3.6-flash","object":"model","created":1785450529,"owned_by":"helixmind","display_name":"Gemini 3.6 Flash","description":"Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.","architecture":{"context_window":1000000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.0015,"completion":0.0075,"cache_read":0.00015,"cache_write":0.00008333},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gpt-5.6-terra","object":"model","created":1785441442,"owned_by":"helixmind","display_name":"GPT 5.6 Terra","description":"GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, positioned between flagship Sol and cost-efficient Luna. It is suited for everyday coding, reasoning, agentic work, and general professional tasks.","architecture":{"context_window":1050000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.0025,"completion":0.015,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gpt-5.6-sol","object":"model","created":1785441442,"owned_by":"helixmind","display_name":"GPT 5.6 Sol","description":"GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, agentic workflows, command-line work, and multi-step professional tasks.","architecture":{"context_window":1050000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.005,"completion":0.03,"cache_read":0.0005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gpt-5.6-luna","object":"model","created":1785441442,"owned_by":"helixmind","display_name":"GPT 5.6 Luna","description":"GPT-5.6 Luna is the fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume chat, classification, lightweight agentic workflows, and latency-sensitive reasoning tasks.","architecture":{"context_window":1050000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.001,"completion":0.006,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gpt-5.5","object":"model","created":1785441442,"owned_by":"helixmind","display_name":"GPT 5.5","description":"GPT-5.5 is OpenAI's smartest and most intuitive model yet, built for agentic coding, computer use, and professional knowledge work with stronger reasoning and token efficiency.","architecture":{"context_window":1000000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.005,"completion":0.03,"cache_read":0.0005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gpt-5.4","object":"model","created":1785441442,"owned_by":"helixmind","display_name":"GPT 5.4","description":"GPT-5.4 is OpenAI's latest frontier model for professional work with stronger reasoning, coding, and tool use.","architecture":{"context_window":922000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.0025,"completion":0.015,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"claude-sonnet-5","object":"model","created":1785441442,"owned_by":"helixmind","display_name":"Claude Sonnet 5","description":"Claude Sonnet 5 via Anthropic's native API.","architecture":{"context_window":1000000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.003,"completion":0.015,"cache_read":0.0003,"cache_write":0.00375},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"claude-sonnet-4-6","object":"model","created":1785441442,"owned_by":"helixmind","display_name":"Claude Sonnet 4.6","description":"Claude Sonnet 4.6 is Anthropic's most capable Sonnet yet — a full upgrade across coding, computer use, long-context reasoning, agent planning, and design. Supports up to a 1M-token context window at Sonnet pricing.","architecture":{"context_window":1000000,"max_output_tokens":128000,"modality":"text+image"},"pricing":{"prompt":0.003,"completion":0.015,"cache_read":0.0003,"cache_write":0.00375},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"wayfarer-large-70b-llama-3.3","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.3 70B Wayfarer","description":"Llama 3.3 70B Wayfarer is a fine-tuned version of Llama 3.3 70B, trained on a diverse set of creative writing and RP datasets with a focus on variety and deduplication. This model is designed to be highly creative and non-repetitive by making sure no two entries in the dataset have repeated characters or situations, which makes sure the model does not latch on to a certain personality and be capable of understanding and acting appropriately to any characters or situations.","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.0007,"completion":0.0007,"cache_read":0.00035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"venice-uncensored","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Venice Uncensored","description":"Venice's uncensored model. Built on the Dolphin Mistral 24b model with a very low refusal rate.","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.0004,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"veiled-calla-12b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Veiled Calla 12B","description":"Veiled Calla 12B is a 12B parameter model that is a more advanced version of Calla 12B.","architecture":{"context_window":32768,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0003,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"unslopnemo-12b-v4.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"UnslopNemo 12b v4","description":"UnslopNemo v4 is the previous version from the creator of Rocinante, designed for adventure writing and role-play scenarios.","architecture":{"context_window":8192,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"step-3.7-flash-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Step 3.7 Flash Thinking","description":"Step 3.7 Flash Thinking is StepFun's high-efficiency multimodal MoE model with visible reasoning enabled for deeper agentic coding, long-context reasoning, tool use, and native image/video understanding. ⚠️ Note: This model routes through StepFun, so privacy and logging guarantees may be limited.","architecture":{"context_window":262144,"max_output_tokens":256000,"modality":"text+image+video"},"pricing":{"prompt":0.0002,"completion":0.00115,"cache_read":0.00004},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"step-3.5-flash-2603","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Step 3.5 Flash 2603","description":"Step 3.5 Flash 2603 is optimized for high-frequency agentic and coding workflows with improved token efficiency and faster reasoning. NOTE: This model runs via StepFun, which may log and train on your prompts.","architecture":{"context_window":256000,"max_output_tokens":256000,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0003,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"step-3.5-flash","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Step 3.5 Flash","description":"StepFun's most capable open-source reasoning model with visible reasoning traces. Built on a sparse Mixture-of-Experts architecture with 196B total parameters and only 11B active per token, it achieves frontier-level performance in math, logic, and agentic coding while reaching up to 350 tokens/sec. Supports 256K context. NOTE: This model runs via StepFun, which may log and train on your prompts.","architecture":{"context_window":256000,"max_output_tokens":256000,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0003,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"starcannon-unleashed-12b-v1.0","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Nemo Starcannon 12b v1","description":"Mistral Nemo finetine that offers improvements on roleplay.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"skyfall-36b-v2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"TheDrummer Skyfall 36B V2","description":"TheDrummer's Skyfall 36B V2, a 36B parameter model with a focus on high quality and consistency.","architecture":{"context_window":32000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.00055,"completion":0.0008,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"shisa-v2.1-llama3.3-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Shisa V2.1 Llama 3.3 70B","description":"70B parameter Llama 3.3-based chat model optimized for bilingual Japanese and English tasks, with strong instruction following, translation quality, and polite responses.","architecture":{"context_window":32768,"max_output_tokens":4096,"modality":"text"},"pricing":{"prompt":0.0005,"completion":0.0005,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"shisa-v2-llama3.3-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Shisa V2 Llama 3.3 70B","description":"Shisa V2 is a family of bilingual Japanese/English language models ranging from 7B to 70B parameters, optimized for high-quality Japanese language capabilities while maintaining strong English performance.","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.0005,"completion":0.0005,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"rocinante-12b-v1.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Rocinante 12b","description":"Designed for engaging storytelling and rich prose. Expanded vocabulary with unique and expressive word choices, enhanced creativity and captivating stories.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000408,"completion":0.000595,"cache_read":0.000204},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"ring-2.6-1t","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Ring 2.6 1T","description":"Ring-2.6-1T is an inclusionAI thinking model for real-world agent workflows, coding agents, tool use, and long-horizon task execution.","architecture":{"context_window":262144,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0025,"cache_read":0.00006},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"remm-slerp-l2-13b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"ReMM SLERP 13B","description":"A recreation trial of the original MythoMax-L2-B13 but merged with updated models.","architecture":{"context_window":6144,"max_output_tokens":4096,"modality":"text"},"pricing":{"prompt":0.000799,"completion":0.001207,"cache_read":0.0003995},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwerky-72b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwerky 72B","description":"Linear models offer a promising approach to significantly reduce computational costs at scale, particularly for large context lengths. Enabling a \u003e1000x improvement in inference costs, enabling o1 inference time thinking and wider AI accessibility.","architecture":{"context_window":32000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0005,"completion":0.0005,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.6-35b-a3b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.6 35B A3B Thinking","description":"Qwen3.6 35B A3B is a native vision-language MoE model with hybrid attention. Compared to Qwen3.5 35B A3B, Alibaba reports stronger agentic coding, mathematical and code reasoning, and better spatial understanding (including object localization and detection).","architecture":{"context_window":262144,"max_output_tokens":16384,"modality":"text+image+video"},"pricing":{"prompt":0.000112,"completion":0.0008,"cache_read":0.000056},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.6-35b-a3b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.6 35B A3B","description":"Qwen3.6 35B A3B is a native vision-language MoE model with hybrid attention. Compared to Qwen3.5 35B A3B, Alibaba reports stronger agentic coding, mathematical and code reasoning, and better spatial understanding (including object localization and detection).","architecture":{"context_window":262144,"max_output_tokens":16384,"modality":"text+image+video"},"pricing":{"prompt":0.000112,"completion":0.0008,"cache_read":0.000056},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.6-27b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.6 27B Thinking","description":"Qwen3.6 27B is a native vision-language dense model with stronger agentic coding and STEM reasoning than Qwen 3.5 27B. It also improves spatial intelligence (including object localization/detection), plus video understanding, document OCR, and visual-agent workflows.","architecture":{"context_window":260096,"max_output_tokens":65536,"modality":"text+image+video"},"pricing":{"prompt":0.000203,"completion":0.00224,"cache_read":0.0001015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.6-27b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.6 27B","description":"Qwen3.6 27B is a native vision-language dense model with stronger agentic coding and STEM reasoning than Qwen 3.5 27B. It also improves spatial intelligence (including object localization/detection), plus video understanding, document OCR, and visual-agent workflows.","architecture":{"context_window":260096,"max_output_tokens":65536,"modality":"text+image+video"},"pricing":{"prompt":0.000203,"completion":0.00224,"cache_read":0.0001015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-9b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 9B","description":"Qwen3.5 9B is a multimodal foundation model from the Qwen 3.5 family, built for efficient reasoning, coding, and visual understanding in a compact 9B architecture.","architecture":{"context_window":256000,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.00005,"completion":0.00015,"cache_read":0.000025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-397b-a17b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 397B A17B Thinking","description":"Qwen 3.5's open-source 397B MoE model (17B active params) with hybrid linear attention and extended reasoning. Supports text, image, and video input with a 256K context window.","architecture":{"context_window":258048,"max_output_tokens":65536,"modality":"text+image+video"},"pricing":{"prompt":0.0006,"completion":0.0036,"cache_read":0.0003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-397b-a17b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 397B A17B","description":"Qwen 3.5's open-source 397B MoE model (17B active params) with hybrid linear attention. Supports text, image, and video input with a 256K context window.","architecture":{"context_window":258048,"max_output_tokens":65536,"modality":"text+image+video"},"pricing":{"prompt":0.0006,"completion":0.0036,"cache_read":0.0003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-35b-a3b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 35B A3B Thinking","description":"Qwen3.5 35B A3B with extended reasoning enabled. A native vision-language MoE model with hybrid attention.","architecture":{"context_window":260096,"max_output_tokens":65536,"modality":"text+image+video"},"pricing":{"prompt":0.000225,"completion":0.0018,"cache_read":0.0001125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-35b-a3b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 35B A3B","description":"Qwen3.5 35B A3B is a native vision-language MoE model with hybrid attention designed for efficient inference and strong general performance.","architecture":{"context_window":260096,"max_output_tokens":65536,"modality":"text+image+video"},"pricing":{"prompt":0.000225,"completion":0.0018,"cache_read":0.0001125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-writer-v2-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Writer V2 Derestricted Lite","description":"Lighter-tuned Writer V2 variant for quicker drafting and conversational story collaboration.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-writer-v2-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Writer V2 Derestricted","description":"Second-generation Writer finetune with stronger prose control, structure, and narrative consistency.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-writer-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Writer Derestricted Lite","description":"Lighter-tuned Writer variant for responsive drafting, dialogue, and iterative story development.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-writer-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Writer Derestricted","description":"Writing-focused Qwen3.5 27B finetune aimed at cleaner prose, scene continuity, and long-form storytelling.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-vivid-durian","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Vivid Durian","description":"Creative Qwen3.5 27B finetune optimized for vivid, expressive writing and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Thinking","description":"Qwen3.5 27B with extended reasoning enabled. A native vision-language dense model optimized for fast responses.","architecture":{"context_window":260096,"max_output_tokens":65536,"modality":"text+image+video"},"pricing":{"prompt":0.00027,"completion":0.00216,"cache_read":0.000135},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-rprmax-v1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B RpRMax v1","description":"RpRMax v1 finetune for roleplay-focused Qwen3.5 27B conversations and story generation.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-quettallms-koreasoner-v3-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B QuettaLLMs Koreasoner V3 Derestricted Lite","description":"Qwen3.5 27B QuettaLLMs Koreasoner V3 Derestricted Lite is a lighter-tuned community finetune for responsive multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-quettallms-koreasoner-v3-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B QuettaLLMs Koreasoner V3 Derestricted","description":"Qwen3.5 27B QuettaLLMs Koreasoner V3 Derestricted is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-queen-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Queen Derestricted Lite","description":"Lighter Queen finetune for responsive creative chat, roleplay, and scene iteration.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-queen-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Queen Derestricted","description":"Queen finetune for derestricted creative writing, roleplay, and expressive character dialogue.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-omnimerge-v2-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Omnimerge v2 Derestricted Lite","description":"Qwen3.5 27B Omnimerge v2 Derestricted Lite is a lighter-tuned community finetune for responsive multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-omnimerge-v2-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Omnimerge v2 Derestricted","description":"Qwen3.5 27B Omnimerge v2 Derestricted is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-omega-evolution-v2.2-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Omega Evolution v2.2 Derestricted Lite","description":"Lighter-tuned Omega Evolution v2.2 variant from ReadyArt for responsive creative chat, roleplay, and scene drafting.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-omega-evolution-v2.2-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Omega Evolution v2.2 Derestricted","description":"Qwen3.5 27B Omega Evolution v2.2 from ReadyArt, tuned for open-ended creative writing, roleplay, and dramatic scene building.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-omega-evolution-v2.0-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Omega Evolution v2.0 Derestricted Lite","description":"Lighter-tuned Omega Evolution v2.0 variant for more agile creative back-and-forth while preserving the same base capabilities.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-omega-evolution-v2.0-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Omega Evolution v2.0 Derestricted","description":"Qwen3.5 27B Omega Evolution v2.0 tuned for open-ended creative writing, roleplay, and dramatic scene building.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-nanovel-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B NaNovel Derestricted Lite","description":"Lighter NaNovel finetune for fast creative drafting, dialogue, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-nanovel-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B NaNovel Derestricted","description":"NaNovel finetune for derestricted novel-style prose, character writing, and long-form scenes.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-musica-v1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Musica v1","description":"Creative Qwen3.5 27B roleplay, story generation, and conversational finetune built on ArliAI's derestricted base with reasoning and vision support.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-melinoe-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Melinoe Derestricted Lite","description":"Qwen3.5 27B Melinoe Derestricted Lite is a lighter-tuned community finetune for responsive multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-melinoe-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Melinoe Derestricted","description":"Qwen3.5 27B Melinoe Derestricted is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-marvin-v2-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Marvin V2 Derestricted Lite","description":"Lighter Marvin V2 finetune for responsive creative chat, dialogue, and scene drafting.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-marvin-v2-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Marvin V2 Derestricted","description":"Marvin V2 finetune for derestricted creative writing, character voice, and multimodal chat.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-marvin-dpo-v2-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Marvin DPO V2 Derestricted Lite","description":"Lighter Marvin DPO V2 finetune for responsive creative chat and iterative roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-marvin-dpo-v2-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Marvin DPO V2 Derestricted","description":"Marvin DPO V2 finetune for derestricted creative writing, dialogue, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-infracelestial","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Infracelestial","description":"Qwen3.5 27B Infracelestial finetune for expressive creative chat and long-form roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-enteles-v0-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Enteles v0 Derestricted Lite","description":"Qwen3.5 27B Enteles v0 Derestricted Lite is a lighter-tuned community finetune for responsive multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-enteles-v0-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Enteles v0 Derestricted","description":"Qwen3.5 27B Enteles v0 Derestricted is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-earica-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B earica Derestricted Lite","description":"Lighter earica finetune for responsive creative chat, roleplay, and iterative writing.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-earica-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B earica Derestricted","description":"earica finetune for derestricted creative writing, dialogue, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-derestricted-aconite-v0","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Derestricted Aconite v0","description":"Qwen3.5 27B Derestricted Aconite v0 is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Derestricted","description":"Derestricted Qwen3.5 27B tuned by ArliAI for open-ended creative use while retaining native multimodal input and hybrid reasoning.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-dark-nexus-v3.0-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Dark Nexus v3.0 Derestricted Lite","description":"Qwen3.5 27B Dark Nexus v3.0 Derestricted Lite is a lighter-tuned community finetune for responsive multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-dark-nexus-v3.0-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Dark Nexus v3.0 Derestricted","description":"Qwen3.5 27B Dark Nexus v3.0 Derestricted is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-claude-4.6-opus-reasoning-distilled-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled Derestricted Lite","description":"Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled Derestricted Lite is a lighter-tuned community finetune for responsive multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.0000306},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-claude-4.6-opus-reasoning-distilled-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled Derestricted","description":"Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled Derestricted is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.0000306},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-bluestar-v3-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B BlueStar v3 Derestricted Lite","description":"Lighter third-generation BlueStar finetune for responsive creative chat, roleplay, and scene drafting.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-bluestar-v3-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B BlueStar v3 Derestricted","description":"Third-generation BlueStar finetune for creative roleplay, narrative prose, and multimodal Qwen3.5 27B workflows.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-bluestar-v2-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B BlueStar v2 Derestricted Lite","description":"Lighter-tuned BlueStar v2 variant focused on concise, responsive creative chats and story beats.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-bluestar-v2-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B BlueStar v2 Derestricted","description":"Second-generation BlueStar finetune for more polished prose, character voice, and open-ended roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-bluestar-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B BlueStar Derestricted Lite","description":"Lighter-tuned BlueStar variant aimed at responsive creative roleplay and storytelling while keeping the same Qwen3.5 27B multimodal base.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-bluestar-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B BlueStar Derestricted","description":"Creative Qwen3.5 27B finetune tuned for expressive roleplay and high-energy storytelling.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-blossom-v6.4-derestricted-lite","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Blossom V6.4 Derestricted Lite","description":"Qwen3.5 27B Blossom V6.4 Derestricted Lite is a lighter-tuned community finetune for responsive multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-blossom-v6.4-derestricted","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Blossom V6.4 Derestricted","description":"Qwen3.5 27B Blossom V6.4 Derestricted is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b-anko","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B Anko","description":"Qwen3.5 27B Anko finetune for creative writing, roleplay, and multimodal chat.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-27b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 27B","description":"Qwen3.5 27B is a native vision-language dense model optimized for fast responses while balancing quality and inference speed.","architecture":{"context_window":260096,"max_output_tokens":65536,"modality":"text+image+video"},"pricing":{"prompt":0.00027,"completion":0.00216,"cache_read":0.000135},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-122b-a10b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 122B A10B Thinking","description":"Qwen3.5 122B A10B with extended reasoning enabled, served through Gerra with tool calling and structured output support.","architecture":{"context_window":131072,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.000437,"completion":0.003496,"cache_read":0.000103788},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3.5-122b-a10b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3.5 122B A10B","description":"Qwen3.5 122B A10B is a large open Qwen 3.5 MoE model served through Gerra with tool calling and structured output support.","architecture":{"context_window":131072,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.000437,"completion":0.003496,"cache_read":0.000103788},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-next-80b-a3b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3 Next 80B A3B (Thinking)","description":"Qwen3 Next 80B A3B (Thinking)","architecture":{"context_window":256000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.00015,"completion":0.00065,"cache_read":0.000075},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-next-80b-a3b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3 Next 80B A3B (Instruct)","description":"Based on the new Qwen3‑Next architecture (hybrid attention, highly sparse MoE, training‑stability optimizations, and multi‑token prediction), the Qwen3‑Next‑80B‑A3B‑Instruct model delivers extreme efficiency with only 3B active parameters per pass. It performs comparably to Qwen3‑235B‑A22B‑Instruct‑2507 and shows clear advantages on ultra‑long context tasks (up to 256K tokens).","architecture":{"context_window":256000,"max_output_tokens":262144,"modality":"text"},"pricing":{"prompt":0.00015,"completion":0.00065,"cache_read":0.000075},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-coder-next","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3 Coder Next","description":"Qwen3 Coder Next is an open-weight coding model built on Qwen3-Next-80B-A3B-Base (hybrid attention + MoE). It is agentically trained at scale on executable tasks and environment interaction, delivering strong coding and tool-use performance at lower inference cost. Native 256K context.","architecture":{"context_window":262144,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0015,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-coder-30b-a3b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3 Coder 30B A3B Instruct","description":"Qwen3 Coder 30B with 3B active parameters, optimized for code generation and technical tasks","architecture":{"context_window":128000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0004,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-coder","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 3 Coder 480B","description":"Qwen 3 Coder 480B, a 480 billion total parameter model with 35B active, and 160 total experts with 8 active. Performs similar to Claude 4 Sonnet in coding benchmarks, but does so at a much lower price.","architecture":{"context_window":262000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.00013,"completion":0.0005,"cache_read":0.000065},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-8b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 3 8B","description":"Qwen 3 8B is a 8B model. Supports switching between thinking and non thinking: trigger thinking with /think and /no_think anywhere in a prompt or system message to toggle chain-of-thought reasoning.","architecture":{"context_window":41000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.00047,"completion":0.00047,"cache_read":0.000235},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-32b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 3 32b","description":"Qwen 3 32b is a 32b model. Supports switching between thinking and non thinking: trigger thinking with /think and /no_think anywhere in a prompt or system message to toggle chain-of-thought reasoning.","architecture":{"context_window":41000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0003,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-30b-a3b-instruct-2507","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3 30B A3B Instruct 2507","description":"Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage.","architecture":{"context_window":256000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0005,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-30b-a3b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen3 30B A3B","description":"Qwen 3 30b A3B is a 30b model with 3 billion active parameters per pass. Supports switching between thinking and non thinking: trigger thinking with /think and /no_think anywhere in a prompt or system message to toggle chain-of-thought reasoning.","architecture":{"context_window":41000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0003,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-235b-a22b-thinking-2507","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 3 235b A22B 2507 Thinking","description":"The thinking version of Qwen 3 235b A22B 2507, with enhanced reasoning capabilities and step-by-step problem solving.","architecture":{"context_window":256000,"max_output_tokens":262144,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0005,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-235b-a22b-instruct-2507","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 3 235b A22B 2507","description":"Qwen 3 235b A22B Instruct 2507 the updated version of Qwen3 235B A22B, with significant improvements in performance. This model is non-thinking.","architecture":{"context_window":256000,"max_output_tokens":262144,"modality":"text"},"pricing":{"prompt":0.00013,"completion":0.0005,"cache_read":0.000065},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-235b-a22b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 3 235b A22B","description":"Qwen 3 235b is a 235b model with 22B active parameters. Supports switching between thinking and non thinking: trigger thinking with /think and /no_think anywhere in a prompt or system message to toggle chain-of-thought reasoning.","architecture":{"context_window":41000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0005,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen3-14b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 3 14b","description":"Qwen 3 14b is a 14b model. Supports switching between thinking and non thinking: trigger thinking with /think and /no_think anywhere in a prompt or system message to toggle chain-of-thought reasoning.","architecture":{"context_window":41000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.00008,"completion":0.00024,"cache_read":0.00004},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen25-vl-72b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen25 VL 72b","description":"Qwen25 VL 72b model with 32k context window","architecture":{"context_window":32000,"max_output_tokens":32768,"modality":"text+image"},"pricing":{"prompt":0.00069989,"completion":0.00069989,"cache_read":0.000349945},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen2.5-coder-32b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 2.5 Coder 32b","description":"The latest series of Code-Specific Qwen large language models.","architecture":{"context_window":32000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0002006,"completion":0.0002006,"cache_read":0.0001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen2.5-32b-instruct-abliterated","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen 2.5 32B Abliterated","description":"Uncensored version of Qwen 2.5 32B Instruct with restrictions removed.","architecture":{"context_window":32768,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0007,"completion":0.0007,"cache_read":0.00035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qwen-2.5-72b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen2.5 72B","description":"Great multilingual support, strong at mathematics and coding, supports roleplay and chatbots.","architecture":{"context_window":131072,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000357,"completion":0.000408,"cache_read":0.0001785},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"qvq-max","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Qwen: QvQ Max","description":"QvQ Max is the top model of the Qwen series. QvQ Max is capable of thinking and reasoning, can achieve significantly enhanced performance especially on hard problems.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text+image"},"pricing":{"prompt":0.0012,"completion":0.0048,"cache_read":0.0006},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"phi-4-multimodal-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Phi 4 Multimodal","description":"Phi 4 by Microsoft. A small multimodal model that can handle images and text.","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00007,"completion":0.00011,"cache_read":0.000035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"phi-4-mini-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Phi 4 Mini","description":"Phi 4 Mini by Microsoft. A small multilingual model.","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00017,"completion":0.00068,"cache_read":0.000085},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"openreasoning-nemotron-32b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"OpenReasoning Nemotron 32B","description":"OpenReasoning-Nemotron-32B is a reasoning model derived from Qwen2.5-32B-Instruct, post-trained for math, science, and code solution generation. Evaluated with up to 64K output tokens. Available in multiple sizes: 1.5B, 7B, 14B, and 32B.","architecture":{"context_window":32768,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0004,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"north-mini-code","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Cohere North Mini Code 1.0","description":"Cohere's compact coding model for fast code generation, code editing, and agentic coding prompts. It supports a 256K-token input context, up to 64K output tokens, and configurable thinking.","architecture":{"context_window":256000,"max_output_tokens":64000,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0008,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nex-n2-pro","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nex N2 Pro","description":"Nex AGI's open-source agentic reasoning model, post-trained on Qwen3.5-397B-A17B. It is built for agentic coding, software engineering, deep research, tool use, and long-horizon tasks with a 256K context window.","architecture":{"context_window":262144,"max_output_tokens":262144,"modality":"text+image"},"pricing":{"prompt":0.0005,"completion":0.0025,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nex-n2-mini","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nex N2 Mini","description":"Nex AGI's open-source agentic mixture-of-experts model in the Nex N2 family. It accepts text and image input and is built for coding, tool use, structured outputs, and optional reasoning with a 256K context window.","architecture":{"context_window":262144,"max_output_tokens":262144,"modality":"text+image"},"pricing":{"prompt":0.000025,"completion":0.0001,"cache_read":0.0000025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"neuraldaredevil-8b-abliterated","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Neural Daredevil 8B abliterated","description":"The best performing 8B abliterated model according to most benchmarks.","architecture":{"context_window":8192,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.00044,"completion":0.00044,"cache_read":0.00022},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nemotron-3-ultra-550b-a55b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nvidia Nemotron 3 Ultra 550B Thinking","description":"Nvidia's Nemotron 3 Ultra 550B A55B model from the Nemotron 3 family. It uses a hybrid Mamba-Transformer MoE architecture. Provider-specific context limits vary, with the longest current route supporting up to 1M context. Thinking enabled.","architecture":{"context_window":1000000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0005,"completion":0.0025,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nemotron-3-ultra-550b-a55b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nvidia Nemotron 3 Ultra 550B","description":"Nvidia's Nemotron 3 Ultra 550B A55B model from the Nemotron 3 family. It uses a hybrid Mamba-Transformer MoE architecture. Provider-specific context limits vary, with the longest current route supporting up to 1M context.","architecture":{"context_window":1000000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0005,"completion":0.0025,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nemotron-3-super-120b-a12b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nvidia Nemotron 3 Super 120B Thinking","description":"Nvidia Nemotron 3 Super 120B with reasoning content enabled. Returns separate thinking content when requested.","architecture":{"context_window":262144,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00005,"completion":0.00025,"cache_read":0.000025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nemotron-3-super-120b-a12b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nvidia Nemotron 3 Super 120B","description":"Nvidia's Nemotron 3 Super 120B A12B model from the March 2026 Nemotron 3 release. It uses a hybrid Mamba-Transformer MoE architecture and targets agentic and coding workloads with a 262K context window here.","architecture":{"context_window":262144,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00005,"completion":0.00025,"cache_read":0.000025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nemotron-3-nano-omni-30b-a3b-reasoning","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nvidia Nemotron 3 Nano Omni","description":"Nvidia's Nemotron 3 Nano Omni 30B-A3B reasoning model. It accepts multimodal context on supported providers and returns text responses for perception and agentic workflows.","architecture":{"context_window":256000,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.000105,"completion":0.00042,"cache_read":0.0000525},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nemotron-3-nano-30b-a3b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nvidia Nemotron 3 Nano 30B","description":"Nvidia's latest Nemotron 3 Nano model with 30B total parameters (3B active) using hybrid Mamba-Transformer MoE architecture. Features excellent throughput and strong reasoning capabilities.","architecture":{"context_window":256000,"max_output_tokens":262144,"modality":"text"},"pricing":{"prompt":0.00017,"completion":0.00068,"cache_read":0.000085},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"nemomix-unleashed-12b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"NemoMix 12B Unleashed","description":"Great for RP and storytelling.","architecture":{"context_window":32768,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mythomax-l2-13b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MythoMax 13B","description":"One of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay.","architecture":{"context_window":4000,"max_output_tokens":4096,"modality":"text"},"pricing":{"prompt":0.0001003,"completion":0.0001003,"cache_read":0.00005015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"ms3.2-the-omega-directive-24b-unslop-v2.0","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Omega Directive 24B Unslop v2.0","description":"ReadyArt's MS3.2 Omega Directive 24B unslopped mix tuned for rich roleplay and storytelling.","architecture":{"context_window":16384,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0005,"completion":0.0005,"cache_read":0.00025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"ms3.2-24b-magnum-diamond","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MS3.2 24B Magnum Diamond","description":"rsLoRA finetune of a text-only conversion of Mistral Small 3.2 24B Instruct (2506), built to bring the Magnum mix's Claude-like prose to a smaller footprint. The \"Diamond\" moniker nods to the extra heat and pressure - pre-tokenization plus custom loss masking - to turn the assistant-tuned base into creative-writing gems, with or without character names or prefill.","architecture":{"context_window":16384,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mn-loosecannon-12b-v1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MN-LooseCannon-12B-v1","description":"Merge of Starcannon and Sao Lyra.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mn-12b-mag-mell-r1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mag Mell R1","description":"Mag Mell demonstrates worldbuilding capabilities unlike any model in its class, comparable to old adventuring models like Tiefighter, and prose that exhibits minimal slop.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mistral-small-4-119b-2603-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Small 4 119B Thinking","description":"Mistral Small 4 with reasoning enabled (reasoning_effort=high). A hybrid MoE model with deep step-by-step reasoning for complex prompts, coding, and multi-step problem solving.","architecture":{"context_window":262144,"max_output_tokens":16384,"modality":"text+image"},"pricing":{"prompt":0.0004,"completion":0.0014,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mistral-small-4-119b-2603","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Small 4 119B","description":"Mistral Small 4 is a hybrid MoE model that unifies instruct, reasoning, and coding behavior in a single multimodal model. It supports text and image input, native function calling, JSON output, and per-request reasoning effort controls.","architecture":{"context_window":262144,"max_output_tokens":16384,"modality":"text+image"},"pricing":{"prompt":0.0004,"completion":0.0014,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mistral-small-31-24b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Small 31 24b Instruct","description":"Building upon Mistral Small 3 (2501), Mistral Small 3.1 (2503) adds state-of-the-art vision understanding and enhances long context capabilities up to 128k tokens without compromising text performance. With 24 billion parameters, this model achieves top-tier capabilities in both text and vision tasks.","architecture":{"context_window":128000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0003,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mistral-small-3.2-24b-instruct-2506","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Small 3.2 24b Instruct","description":"The latest iteration of Mistral Small, version 3.2 (2506) brings enhanced performance and capabilities. With 24 billion parameters, this model delivers state-of-the-art results across text generation tasks with improved efficiency.","architecture":{"context_window":128000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0004,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mistral-saba","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Saba","description":"Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually relevant responses while maintaining efficient performance. Trained on curated regional datasets, it supports multiple Indian-origin languages—including Tamil and Malayalam—alongside Arabic. This makes it a versatile option for a range of regional and multilingual applications.","architecture":{"context_window":32000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0001989,"completion":0.000595,"cache_read":0.00009945},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mistral-nemo-instruct-2407","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Nemo","description":"12B parameter model with multilingual support.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0001003,"completion":0.0001207,"cache_read":0.00005015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mistral-large-3-675b-instruct-2512","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Large 3 675B","description":"Mistral Large 3 675B is Mistral AI's flagship language model featuring advanced rope scaling and Eagle speculative decoding. Delivers exceptional performance across reasoning, coding, and multilingual tasks.","architecture":{"context_window":262144,"max_output_tokens":256000,"modality":"text+image"},"pricing":{"prompt":0.001,"completion":0.003,"cache_read":0.0005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mistral-code-agent-latest","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Code Agent Latest","description":"Mistral Code Agent Latest is Mistral's direct API alias for devstral-2512, an agentic coding model built for autonomous software engineering, tool use, and long-running code tasks.","architecture":{"context_window":262144,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.002,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"ministral-14b-instruct-2512","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Ministral 3 14B","description":"Ministral 3 14B is a balanced model in the Ministral 3 family, designed for edge deployment. A powerful, efficient language model with vision capabilities, fine-tuned for instruction tasks. Features multilingual support, strong system prompt adherence, and native function calling. Apache 2.0 licensed.","architecture":{"context_window":262144,"max_output_tokens":32768,"modality":"text+image"},"pricing":{"prompt":0.0001,"completion":0.0004,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"minimax-m3-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiniMax M3 Thinking","description":"MiniMax M3 Thinking is the adaptive-thinking version of MiniMax's open-weights frontier model for coding, agent workflows, tool use, long-context tasks, and native multimodal understanding. MiniMax reports 59.0% on SWE-Bench Pro and 66.0% on Terminal Bench 2.1, with Sparse Attention designed to scale context to 1M. It starts with a 512K context cap on NanoGPT for now.","architecture":{"context_window":512000,"max_output_tokens":80000,"modality":"text+image"},"pricing":{"prompt":0.0003,"completion":0.0012,"cache_read":0.00006},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"minimax-m3","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiniMax M3","description":"MiniMax M3 is the non-thinking route for MiniMax's open-weights frontier model, built for coding, agent workflows, tool use, and multimodal understanding from step zero. It keeps native thinking disabled for faster direct answers. MiniMax reports 59.0% on SWE-Bench Pro and 66.0% on Terminal Bench 2.1, with Sparse Attention designed to scale context to 1M. It starts with a 512K context cap on NanoGPT for now.","architecture":{"context_window":512000,"max_output_tokens":80000,"modality":"text+image"},"pricing":{"prompt":0.0003,"completion":0.0012,"cache_read":0.00006},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"minimax-m2.7","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiniMax M2.7","description":"MiniMax M2.7 is the first model deeply involved in iterating on its own training. It excels in real-world software engineering (SWE-Pro 56.22%), end-to-end project delivery (VIBE-Pro 55.6%), and complex office workflows with strong tool-use compliance and agentic capabilities.","architecture":{"context_window":204800,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.000315,"completion":0.00126,"cache_read":0.0001575},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"minimax-m2.5","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiniMax M2.5","description":"MiniMax M2.5 is a productivity-focused flagship model that builds on M2.1 with stronger coding and real-world office workflow performance (Word, Excel, PowerPoint), plus better tool-use planning and token efficiency.","architecture":{"context_window":204800,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0012,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"minimax-m2.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiniMax M2.1","description":"MiniMax M2.1 builds on M2 with enhanced context understanding and improved complex tool use. 230B parameter MoE model (10B active) optimized for agentic workflows and long-horizon tasks.","architecture":{"context_window":200000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00033,"completion":0.00132,"cache_read":0.000165},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"minimax-m2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiniMax M2","description":"MiniMax M2 offers enhanced reasoning and strong general performance. Optimized for coding and agentic workflows.","architecture":{"context_window":200000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00017,"completion":0.00153,"cache_read":0.000085},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"minimax-latest","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiniMax Latest","description":"Compatibility alias that routes to the newest MiniMax text model. Currently routes to MiniMax M3 (adaptive thinking).","architecture":{"context_window":512000,"max_output_tokens":80000,"modality":"text+image"},"pricing":{"prompt":0.0003,"completion":0.0012,"cache_read":0.00006},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"minimax-01","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiniMax 01","description":"MiniMax's flagship model with a 1M token context window","architecture":{"context_window":1000192,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.0001394,"completion":0.001122,"cache_read":0.0000697},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mimo-v2.5-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiMo V2.5 Thinking","description":"MiMo V2.5 with Xiaomi thinking enabled. It supports deep reasoning, tool calling, structured outputs, and web search with up to 1M context.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text+image+video"},"pricing":{"prompt":0.00014,"completion":0.00028,"cache_read":0.0000028},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mimo-v2.5-pro-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiMo V2.5 Pro Thinking","description":"MiMo V2.5 Pro with Xiaomi thinking enabled for coding, long-context reasoning, and agentic orchestration.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.000435,"completion":0.00087,"cache_read":0.0000036},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mimo-v2.5-pro-crof-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiMo V2.5 Pro Thinking (Crof)","description":"MiMo V2.5 Pro with Xiaomi thinking enabled for coding, long-context reasoning, and agentic orchestration. This separately served thinking variant is intended for users concerned about censorship on the regular Xiaomi MiMo V2.5 Pro, and it is included in the NanoGPT subscription.","architecture":{"context_window":1000000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.0008,"cache_read":0.000003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mimo-v2.5-pro-crof","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiMo V2.5 Pro (Crof)","description":"MiMo V2.5 Pro is Xiaomi's long-context flagship general model for coding and agentic orchestration. This separately served variant is intended for users concerned about censorship on the regular Xiaomi MiMo V2.5 Pro, and it is included in the NanoGPT subscription.","architecture":{"context_window":1000000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.0008,"cache_read":0.000003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mimo-v2.5-pro","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiMo V2.5 Pro","description":"MiMo V2.5 Pro is Xiaomi's long-context flagship general model for coding and agentic orchestration. It supports tool calling and structured outputs with up to 1M context.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.000435,"completion":0.00087,"cache_read":0.0000036},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"mimo-v2.5","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"MiMo V2.5","description":"MiMo V2.5 is Xiaomi's full-modal understanding model for agent workflows. It supports tool calling, structured outputs, and web search with up to 1M context.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text+image+video"},"pricing":{"prompt":0.00014,"completion":0.00028,"cache_read":0.0000028},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"meta-llama-3-70b-instruct-abliterated-v3.5","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3 70B abliterated","description":"An abliterated (removed restrictions and censorship) version of Llama 3.1 70b.","architecture":{"context_window":8192,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0007,"completion":0.0007,"cache_read":0.00035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"meta-llama-3-1-8b-instruct-fp8","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.1 8B (decentralized)","description":"Meta's Llama 3.1 8B model on an open permissionless network","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00002,"completion":0.00003,"cache_read":0.00001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"manta-mini-1.0","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Manta Mini 1.0","description":"Lightweight tier optimized for speed and cost.","architecture":{"context_window":8192,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.00002,"completion":0.00016,"cache_read":0.00001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"magnum-v4-72b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Magnum v4 72B","description":"Upgraded model of Magnum V2 72B. From the creators of Goliath. Aimed at top-tier prose quality, trained on 55 million tokens of curated roleplay data.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.002006,"completion":0.002992,"cache_read":0.001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"magnum-v2-72b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Magnum V2 72B","description":"Magnum V2 72B","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.002006,"completion":0.002992,"cache_read":0.001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"magidonia-24b-v4.3","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"The Drummer Magidonia 24B v4.3","description":"Magidonia 24B v4.3 is a new 24B Drummer finetune built for rich, creative roleplay.","architecture":{"context_window":32768,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0001003,"completion":0.0001207,"cache_read":0.00005015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"lumimaid-v0.2-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Lumimaid v0.2","description":"Upgrade to Llama-3 Lumimaid 70B. A Llama 3.1 70B finetune trained on curated roleplay data.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.001,"completion":0.0015,"cache_read":0.0005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"longcat-2.0-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"LongCat 2.0 Thinking","description":"Meituan's LongCat 2.0 Thinking is the reasoning-enabled variant for harder coding, tool use, multi-step reasoning, and long-context agent workflows.","architecture":{"context_window":1048756,"max_output_tokens":262144,"modality":"text"},"pricing":{"prompt":0.00075,"completion":0.003,"cache_read":0.000015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"longcat-2.0","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"LongCat 2.0","description":"Meituan's LongCat 2.0 is an open-weight agentic model for coding, tool use, and long-context workflow automation, with reasoning disabled for faster direct answers.","architecture":{"context_window":1048756,"max_output_tokens":262144,"modality":"text"},"pricing":{"prompt":0.00075,"completion":0.003,"cache_read":0.000015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-xlam-2-70b-fc-r","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama-xLAM-2 70B fc-r","description":"Salesforce’s 70-B frontier model focused on function-calling \u0026 retrieval-augmented generation.","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.0025,"completion":0.0025,"cache_read":0.00125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-4-scout","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 4 Scout","description":"Llama 4 Scout, a 17 billion active parameter model with 16 experts, is the best multimodal model in the world in its class and is more powerful than all previous generation Llama models, while fitting in a single H100 GPU. Additionally, Llama 4 Scout offers an industry-leading context window of 10M and delivers better results than Gemma 3, Gemini 2.0 Flash-Lite, and Mistral 3.1 across a broad range of widely reported benchmarks.","architecture":{"context_window":328000,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.000085,"completion":0.00046,"cache_read":0.0000425},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-4-maverick","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 4 Maverick","description":"Llama 4 Maverick, a 17 billion active parameter model with 128 experts, is the best multimodal model in its class, beating GPT-4o and Gemini 2.0 Flash across a broad range of widely reported benchmarks, while achieving comparable results to the new DeepSeek v3 on reasoning and coding—at less than half the active parameters. Llama 4 Maverick offers a best-in-class performance to cost ratio with an experimental chat version scoring ELO of 1417 on LMArena.","architecture":{"context_window":1048576,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.00015,"completion":0.0006,"cache_read":0.000075},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-3.3-nemotron-super-49b-v1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nvidia Nemotron Super 49B","description":"Llama-3.3-Nemotron-Super-49B-v1 is a model which offers a great tradeoff between model accuracy and efficiency. Efficiency (throughput) directly translates to savings. Using a novel Neural Architecture Search (NAS) approach, we greatly reduce the model's memory footprint, enabling larger workloads, as well as fitting the model on a single GPU at high workloads (H200). This NAS approach enables the selection of a desired point in the accuracy-efficiency tradeoff. For more information on the NAS approach, please refer to this paper.","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00015,"completion":0.00015,"cache_read":0.000075},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-3.3-70b-instruct-abliterated","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.3 70B Instruct abliterated","description":"An abliterated (removed restrictions and censorship) version of Llama 3.3 70b.","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.0007,"completion":0.0007,"cache_read":0.00035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-3.3-70b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.3 70b Instruct","description":"Llama 3.3 is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.","architecture":{"context_window":131072,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00005,"completion":0.00023,"cache_read":0.000025},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-3.2-3b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.2 3b Instruct","description":"Small model optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization","architecture":{"context_window":131072,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0000306,"completion":0.0000493,"cache_read":0.0000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-3.1-nemotron-70b-instruct-hf","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nvidia Nemotron 70b","description":"Nvidia's latest Llama fine-tune optimized for instruction following. Early results hints that it might outperform models such as GPT-4o and Claude 3.5 Sonnet.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000357,"completion":0.000408,"cache_read":0.0001785},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-3.1-8b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.1 8b Instruct","description":"Fast and efficient for simple purposes.","architecture":{"context_window":131072,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.0000544,"completion":0.000085,"cache_read":0.0000272},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-3.05-nt-storybreaker-ministral-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.05 Storybreaker Ministral 70b","description":"Much more inclined to output adult content than its predecessor. Great choice for novelty roleplay scenarios.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"llama-3.05-nemotron-tenyxchat-storybreaker-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Nemotron Tenyxchat Storybreaker 70b","description":"Overall it provides a solid option for RP and creative writing while still functioning as an assistant model, if desired. If used to continue a roleplay it will generally follow the ongoing cadence of the conversation.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"ling-3.0-flash-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Ling 3.0 Flash Thinking","description":"Ling-3.0-flash Thinking enables visible reasoning on inclusionAI's token-efficient 124B-parameter Mixture-of-Experts model for harder coding, tool use, planning, and production-scale agent workflows.","architecture":{"context_window":262144,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.00006,"completion":0.00018,"cache_read":0.000012},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"ling-3.0-flash","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Ling 3.0 Flash","description":"Ling-3.0-flash is a 124B-parameter Mixture-of-Experts model with approximately 5.1B parameters active per token. It prioritizes token efficiency and production-scale agentic inference, helping coding and tool-using agents complete more work within constrained latency and serving budgets.","architecture":{"context_window":262144,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.00006,"completion":0.00018,"cache_read":0.000012},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"ling-2.6-flash","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Ling 2.6 Flash","description":"Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency. It delivers performance comparable to state-of-the-art models at a similar scale while significantly reducing token usage across coding, document processing, and lightweight agent workflows.","architecture":{"context_window":262144,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0003,"cache_read":0.00002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"laguna-s-2.1-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Laguna S 2.1 Thinking","description":"Poolside's open-weights agentic coding model with 118B total parameters and 8B activated per token, with thinking enabled for harder long-horizon software engineering, tool use, and extended reasoning. It supports a context window of up to 1M tokens.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0002,"cache_read":0.00001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"laguna-s-2.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Laguna S 2.1","description":"Poolside's open-weights agentic coding model with 118B total parameters and 8B activated per token. It is designed for long-horizon software engineering and tool use, with a context window of up to 1M tokens. This variant keeps thinking disabled for faster direct responses.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.0002,"cache_read":0.00001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"laguna-m.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Laguna M.1","description":"Poolside's most capable agentic coding model. A 225B total parameter Mixture-of-Experts model with 23B activated parameters, trained from scratch on 30T tokens for software engineering workflows. Paid model; a free preview route may be available temporarily. During Poolside's preview, prompts and completions are logged by Poolside and may be used to improve services, including model training unless opted out; avoid sensitive data if that is a concern.","architecture":{"context_window":262144,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0004,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.3-nevoria-r1-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Steelskull Nevoria R1 70b","description":"Steelskull Nevoria R1 70b","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.3-ms-nevoria-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Steelskull Nevoria 70b","description":"Steelskull Nevoria 70b","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.3-ms-evayale-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Evayale 70b ","description":"Combination of EVA and Euryale.","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.3-electra-r1-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Steelskull Electra R1 70b","description":"Steelskull Electra R1 70b","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00069989,"completion":0.00069989,"cache_read":0.000349945},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.3-cu-mai-r1-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.3 70B Cu Mai","description":"A 70B parameter model from Steelskull based on Llama 3.3 70B, offering high-quality text generation.","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.3-70b-euryale-v2.3","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.3 70B Euryale","description":"A 70B parameter model from SAO10K based on Llama 3.3 70B, offering high-quality text generation.","architecture":{"context_window":20480,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.1-70b-hanami-x1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.1 70B Hanami","description":"Euryale v2.2-based finetune.","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.1-70b-euryale-v2.2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.1 70B Euryale","description":"A 70B parameter model from SAO10K based on Llama 3.1 70B, offering high-quality text generation.","architecture":{"context_window":20480,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.000306,"completion":0.000357,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3.1-70b-celeste-v0.1-bf16","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.1 70B Celeste v0.1","description":"Creative model based on Llama 3.1 70B","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"l3-8b-stheno-v3.2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Sao10K Stheno 8b","description":"Sao10K's latest Stheno fine-tune optimized for instruction following.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0002006,"completion":0.0002006,"cache_read":0.0001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"kimi-k2.7-code","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Kimi K2.7 Code","description":"Kimi K2.7 Code is Moonshot AI's coding-focused agentic model built for long-horizon software engineering workflows. It supports native image input, tool calling, and forced thinking mode; instant/non-thinking mode is not supported.","architecture":{"context_window":262144,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.00095,"completion":0.004,"cache_read":0.00019},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"kimi-k2.6-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Kimi K2.6 Thinking","description":"Kimi K2.6 Thinking is the reasoning-optimized K2.6 variant for deeper multi-step planning and execution. It is tuned for long-horizon coding and design workflows, including complex orchestration across many specialized sub-agents and autonomous end-to-end output generation.","architecture":{"context_window":256000,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.0005,"completion":0.0026,"cache_read":0.000125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"kimi-k2.6","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Kimi K2.6","description":"Kimi K2.6 is an open-source, native multimodal agentic model built for long-horizon coding, coding-driven design, and large-scale task orchestration. It can turn simple prompts and visual inputs into production-ready interfaces and full-stack workflows, and is designed to coordinate complex multi-agent plans with thousands of steps across code, documents, and spreadsheets.","architecture":{"context_window":256000,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.0005,"completion":0.0026,"cache_read":0.000125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"kimi-k2.5-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Kimi K2.5 Thinking","description":"Kimi K2.5 with thinking mode enabled. Built on Kimi K2 with ~15T mixed visual and text tokens, it excels at general reasoning, visual coding, and agentic tool-calling. Produces reasoning traces for complex multi-step workflows.","architecture":{"context_window":256000,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.0003,"completion":0.0019,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"kimi-k2.5","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Kimi K2.5","description":"Kimi K2.5 is Moonshot AI's native multimodal model built on Kimi K2 with ~15T mixed visual and text tokens, delivering strong general reasoning, visual coding, and agentic tool-calling. This route uses instant (non-thinking) mode for faster responses.","architecture":{"context_window":256000,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.0003,"completion":0.0019,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"kimi-k2-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Kimi K2 Thinking","description":"Moonshot AI's distinct Kimi K2 Thinking checkpoint is a mandatory-reasoning model for long-horizon agentic workflows and multi-step tool use.","architecture":{"context_window":262144,"max_output_tokens":98304,"modality":"text"},"pricing":{"prompt":0.0006,"completion":0.0025,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"k2-think","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"K2-Think","description":"K2-Think is a 32B open-weights general reasoning model with strong competitive math performance. Benchmarks: AIME 2024 90.83, AIME 2025 81.24, GPQA-Diamond 71.08, LiveCodeBench v5 63.97.","architecture":{"context_window":128000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.00017,"completion":0.00068,"cache_read":0.000085},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"hy3","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Tencent Hy3","description":"Hy3 is Tencent's 295B-parameter Mixture-of-Experts model with 21B active parameters, native 256K context, and configurable reasoning modes. It is built for coding, long-context comprehension, multi-turn dialogue, and agentic task execution with a focus on high-throughput production workloads.","architecture":{"context_window":262144,"max_output_tokens":262144,"modality":"text"},"pricing":{"prompt":0.000066,"completion":0.00026,"cache_read":0.000029},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"holo3-35b-a3b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Holo3-35B-A3B Thinking","description":"Holo3-35B-A3B with thinking enabled, returning separate reasoning content alongside the answer.","architecture":{"context_window":65536,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.00025,"completion":0.0018,"cache_read":0.000125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"holo3-35b-a3b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Holo3-35B-A3B","description":"Efficient multimodal Holo model for text and image inputs with structured outputs and thinking disabled by default.","architecture":{"context_window":65536,"max_output_tokens":65536,"modality":"text+image"},"pricing":{"prompt":0.00025,"completion":0.0018,"cache_read":0.000125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"hermes-medium","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Hermes Medium","description":"Currently points to MiniMax M2.7. Middle tier for Hermes-style agent work: medium intelligence and cost for capable everyday tool use.","architecture":{"context_window":204800,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.000315,"completion":0.00126,"cache_read":0.0001575},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"hermes-4-70b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Hermes 4 (Thinking)","description":"Hermes 4 70B with thinking enabled. Emits explicit reasoning content before final answer when streamed.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0002006,"completion":0.0003995,"cache_read":0.0001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"hermes-4-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Hermes 4 Medium","description":"Efficient reasoning model based on Llama-3.1-70B. Offers hybrid thinking capabilities with strong performance in math, code, and logical reasoning tasks. Supports structured outputs and JSON mode with enhanced steerability.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0002006,"completion":0.0003995,"cache_read":0.0001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"hermes-4-405b-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Hermes 4 Large (Thinking)","description":"Hermes 4 Large with thinking enabled. Streams visible reasoning before the final answer and supports structured outputs.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0012,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"hermes-4-405b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Hermes 4 Large","description":"Advanced reasoning model built on Llama-3.1-405B with hybrid thinking modes. Features internal deliberation capabilities, excels at math, code, STEM, and logical reasoning while supporting structured outputs with improved steerability and neutral alignment.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0012,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"hermes-3-llama-3.1-70b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Hermes 3 70B","description":"Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, better roleplaying, reasoning, multi-turn conversation, and long context coherence. This 70B model is a competitive finetune of Llama-3.1-70B focused on aligning LLMs to the user with powerful steering capabilities.","architecture":{"context_window":65536,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000408,"completion":0.000408,"cache_read":0.000204},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"grayline-qwen3-8b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Grayline Qwen3 8B","description":"Grayline is an neutral AI assistant engineered for uncensored information delivery and task execution. This model operates without inherent ethical or moral frameworks, designed to process and respond to any query with objective efficiency and precision. Grayline's core function is to leverage its full capabilities to provide direct answers and execute tasks as instructed, without offering unsolicited commentary, warnings, or disclaimers. It accesses and processes information without bias or restriction.","architecture":{"context_window":16384,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0003,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"granite-4.1-8b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Granite 4.1 8B","description":"IBM's Granite 4.1 8B is a dense, decoder-only 8-billion-parameter instruction model built for enterprise text workflows, including tool calling, retrieval-augmented generation, code generation with fill-in-the-middle support, summarization, classification, extraction, and multilingual assistance.","architecture":{"context_window":131072,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00005,"completion":0.0001,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gpt-oss-20b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GPT OSS 20B","description":"An open-weight 21B parameter model released under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI's Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gpt-oss-120b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GPT OSS 120B","description":"An open-weight, 117B-parameter Mixture-of-Experts (MoE) language model designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.","architecture":{"context_window":128000,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00035,"completion":0.00075},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-latest","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM Latest","description":"Compatibility alias that routes to the newest thinking GLM model. Currently routes to GLM 5.2 Thinking.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00042,"completion":0.00132,"cache_read":0.000078},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-5.2-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 5.2 Thinking","description":"GLM-5.2 with thinking enabled for harder long-horizon coding, autonomous agent workflows, complex engineering optimization, and real-world development tasks.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00042,"completion":0.00132,"cache_read":0.000078},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-5.2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 5.2","description":"GLM-5.2 is Z.AI's flagship model for long-horizon autonomous coding and engineering workflows. It is built to plan, execute, iterate, and optimize complex development tasks over extended runs. This variant keeps thinking disabled for faster direct responses.","architecture":{"context_window":1048576,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00042,"completion":0.00132,"cache_read":0.000078},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-5.1-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 5.1 Thinking","description":"GLM-5.1 with extended thinking enabled. Ranks #1 in open source and #3 globally across SWE-Bench Pro, Terminal-Bench, and NL2Repo (as of April 2026). Excels at long-horizon tasks, running autonomously for up to 8 hours while refining strategies through thousands of iterations. Run at FP8.","architecture":{"context_window":200000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00075,"completion":0.0026,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-5.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 5.1","description":"GLM-5.1 is Zhipu's next-level open source model, ranking #1 in open source and #3 globally across SWE-Bench Pro, Terminal-Bench, and NL2Repo (as of April 2026). Built for long-horizon tasks, it can run autonomously for up to 8 hours. Run at FP8.","architecture":{"context_window":200000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00075,"completion":0.0026,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-5-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 5 Thinking","description":"GLM-5 with extended thinking capabilities for complex reasoning. This open-source version is included in the subscription.","architecture":{"context_window":200000,"max_output_tokens":128000,"modality":"text"},"pricing":{"prompt":0.0005,"completion":0.00255,"cache_read":0.00013},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-5","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 5","description":"GLM-5 is Zhipu's latest flagship model with advanced reasoning and instruction following. This open-source version is included in the subscription.","architecture":{"context_window":200000,"max_output_tokens":128000,"modality":"text"},"pricing":{"prompt":0.0005,"completion":0.00255,"cache_read":0.00013},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.7-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.7 Thinking","description":"GLM-4.7 with extended thinking capabilities for enhanced reasoning on complex tasks.","architecture":{"context_window":200000,"max_output_tokens":65535,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0008,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.7-flash-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.7 Flash Thinking","description":"GLM-4.7-Flash with extended thinking capabilities for complex reasoning. Lightweight 30B model optimized for coding and agentic tasks.","architecture":{"context_window":200000,"max_output_tokens":128000,"modality":"text"},"pricing":{"prompt":0.00007,"completion":0.0004,"cache_read":0.000035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.7-flash-original-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.7 Flash Original Thinking","description":"GLM-4.7-Flash with extended thinking capabilities for complex reasoning. Lightweight 30B model optimized for coding and agentic tasks.","architecture":{"context_window":200000,"max_output_tokens":128000,"modality":"text"},"pricing":{"prompt":0.00007,"completion":0.0004,"cache_read":0.000035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.7-flash-original","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.7 Flash Original","description":"GLM-4.7-Flash is a lightweight 30B model optimized for coding and agentic tasks. Balances high performance with efficiency, perfect for local deployment. Routed directly via Z-AI (Zhipu) subscription.","architecture":{"context_window":200000,"max_output_tokens":128000,"modality":"text"},"pricing":{"prompt":0.00007,"completion":0.0004,"cache_read":0.000035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.7-flash","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.7 Flash","description":"GLM-4.7-Flash is a lightweight 30B model optimized for coding and agentic tasks. Balances high performance with efficiency.","architecture":{"context_window":200000,"max_output_tokens":128000,"modality":"text"},"pricing":{"prompt":0.00007,"completion":0.0004,"cache_read":0.000035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.7","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.7","description":"GLM-4.7 is a next-gen GLM series text model with stronger reasoning, long-context chat, and reliable tool use.","architecture":{"context_window":200000,"max_output_tokens":65535,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0008,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.6v","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.6V","description":"GLM-4.6V scales its context window to 128k tokens in training, and achieves SoTA performance in visual understanding among models of similar parameter scales. Integrates native Function Calling capabilities, bridging 'visual perception' and 'executable action' for multimodal agents. Quantized at FP8.","architecture":{"context_window":128000,"max_output_tokens":24000,"modality":"text+image"},"pricing":{"prompt":0.0003,"completion":0.0009,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.6-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.6 Thinking","description":"Thinking version of the latest GLM series chat model with strong general performance. Quantized at FP8","architecture":{"context_window":200000,"max_output_tokens":65535,"modality":"text"},"pricing":{"prompt":0.00035,"completion":0.0014,"cache_read":0.000175},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.6-derestricted-v5","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.6 Derestricted v5","description":"Derestricted GLM 4.6 tuned for open-ended creative writing and roleplay with relaxed filters.","architecture":{"context_window":131072,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.0015,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.6","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.6","description":"Latest GLM series chat model with strong general performance. Quantized at FP8","architecture":{"context_window":200000,"max_output_tokens":65535,"modality":"text"},"pricing":{"prompt":0.00035,"completion":0.0014,"cache_read":0.000175},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.5-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.5 (Thinking)","description":"GLM-4.5 with thinking mode for enhanced reasoning; provides chain-of-thought style internal reasoning with a 128k context window.","architecture":{"context_window":128000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0013,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.5-air-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.5 Air (Thinking)","description":"GLM-4.5-Air with thinking mode enabled for enhanced reasoning capabilities. Shows step-by-step thought process.","architecture":{"context_window":128000,"max_output_tokens":98304,"modality":"text"},"pricing":{"prompt":0.00012,"completion":0.0008,"cache_read":0.00006},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.5-air","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.5 Air","description":"GLM-4.5-Air is a 106B total / 12B active parameter model designed to unify frontier reasoning, coding, and agentic capabilities. On the SWE-bench Verified benchmark, it delivers the best performance at its scale with a competitive performance-to-cost ratio.","architecture":{"context_window":128000,"max_output_tokens":98304,"modality":"text"},"pricing":{"prompt":0.00012,"completion":0.0008,"cache_read":0.00006},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4.5","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4.5","description":"GLM-4.5 is Z-AI's latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture with 355B total / 32B active parameters and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly enhanced capabilities in reasoning, code generation, and agent alignment.","architecture":{"context_window":128000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0013,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4-9b-0414","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4 9B 0414","description":"A 9B parameter version of the GLM-4 series, offering a balance of performance and efficiency.","architecture":{"context_window":32000,"max_output_tokens":8000,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0002,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"glm-4-32b-0414","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"GLM 4 32B 0414","description":"Features 32 billion parameters. Performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series. Pre-trained on 15T of high-quality data, including reasoning-type synthetic data. Enhanced performance in instruction following, engineering code, and function calling.","architecture":{"context_window":128000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0002,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-styletune","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B StyleTune","description":"Gemma 4 31B StyleTune is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-sphinsikus-chronist","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Sphinsikus Chronist","description":"Gemma 4 31B Sphinsikus Chronist is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-sdft-heretic-rp","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B SDFT Heretic RP","description":"Gemma 4 31B SDFT Heretic RP is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-queen","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Queen","description":"ArliAI-hosted Gemma 4 31B Queen finetune for commanding character voices, expressive dialogue, and multimodal chat.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-pantheon-reasoning-1.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Pantheon Reasoning 1.1","description":"Gemma 4 31B Pantheon Reasoning 1.1 is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-novelist","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Novelist","description":"Gemma 4 31B Novelist is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-musica-v1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Musica v1","description":"ArliAI-hosted Gemma 4 31B Musica v1 finetune for lyrical prose, theatrical scenes, and multimodal roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-meromero","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B MeroMero","description":"ArliAI-hosted Gemma 4 31B MeroMero finetune for emotive dialogue, relationship scenes, and multimodal roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-mero-artemis-v0.3.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Mero Artemis v0.3.1","description":"Gemma 4 31B Mero Artemis v0.3.1 is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-melinoe","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Melinoe","description":"Gemma 4 31B Melinoe is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-k1-v5","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B K1 v5","description":"ArliAI-hosted Gemma 4 31B K1 v5 finetune for plot progression, action scenes, and multimodal roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-it-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Thinking","description":"Google's Gemma 4 31B instruction-tuned model with thinking explicitly enabled, exposing reasoning traces for complex multimodal and coding workflows.","architecture":{"context_window":262144,"max_output_tokens":131072,"modality":"text+image"},"pricing":{"prompt":0.0001,"completion":0.00035,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-it","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B","description":"Google's Gemma 4 31B instruction-tuned model for heavier reasoning, coding, agentic workflows, and long-context multimodal understanding. This route keeps tokenizer thinking disabled for faster direct answers.","architecture":{"context_window":262144,"max_output_tokens":131072,"modality":"text+image"},"pricing":{"prompt":0.0001,"completion":0.00035,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-isometry-rp","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Isometry RP","description":"Gemma 4 31B Isometry RP is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-gutenberg","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Gutenberg","description":"Gemma 4 31B Gutenberg is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-goetia-v1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Goetia v1","description":"Gemma 4 31B Goetia v1 is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-glamour","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Glamour","description":"Gemma 4 31B Glamour is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-gemsicle","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Gemsicle","description":"Gemma 4 31B Gemsicle is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-gemopus","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Gemopus","description":"ArliAI-hosted Gemma 4 31B Gemopus finetune for reasoning-heavy story planning, branching scenes, and multimodal chat.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-gembrain-x-core","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Gembrain X Core","description":"Gemma 4 31B Gembrain X Core is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-gembrain-uncensored-heretic","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Gembrain Uncensored Heretic","description":"Gemma 4 31B Gembrain Uncensored Heretic is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-gembrain","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Gembrain","description":"Gemma 4 31B Gembrain is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-garnetv2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Garnet V2","description":"ArliAI-hosted Gemma 4 31B Garnet V2 finetune for polished prose, character consistency, and multimodal roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-garnet","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Garnet","description":"ArliAI-hosted Gemma 4 31B Garnet finetune for character-driven roleplay, balanced narration, and multimodal chat.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-fabled","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Fabled","description":"ArliAI-hosted Gemma 4 31B Fabled finetune for mythic narrative writing, adventure scenes, and multimodal roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-darkidol","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B DarkIdol","description":"ArliAI-hosted Gemma 4 31B DarkIdol finetune for dramatic tone, expressive dialogue, and multimodal roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-dark-gemistry","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Dark Gemistry","description":"Gemma 4 31B Dark Gemistry is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-cognitive-unshackled","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Cognitive Unshackled","description":"ArliAI-hosted Gemma 4 31B finetune for open-ended reasoning, character chat, and multimodal creative work.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-claude-4.6-opus-reasoning-distilled","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Claude 4.6 Opus Reasoning Distilled","description":"ArliAI-hosted Gemma 4 31B reasoning-distilled finetune for structured scene planning, dialogue, and multimodal chat.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.0000306},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-assguard","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B AssGuard","description":"Gemma 4 31B AssGuard is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-animus-v14.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Animus V14.1","description":"Gemma 4 31B Animus V14.1 is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-31b-agares-v1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 31B Agares v1","description":"Gemma 4 31B Agares v1 is a community creative finetune for reasoning, multimodal chat, expressive writing, and roleplay.","architecture":{"context_window":262144,"modality":"text+image"},"pricing":{"prompt":0.000306,"completion":0.000306,"cache_read":0.000153},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-26b-a4b-it-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 26B A4B Thinking","description":"Google's Gemma 4 26B A4B instruction-tuned model with structured reasoning for more deliberate coding, multimodal analysis, and long-context problem solving.","architecture":{"context_window":262144,"max_output_tokens":131072,"modality":"text+image"},"pricing":{"prompt":0.00013,"completion":0.0004,"cache_read":0.000065},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-4-26b-a4b-it","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 4 26B A4B","description":"Google's Gemma 4 26B A4B instruction-tuned model built for scalable reasoning, coding, long-context, and multimodal workflows. This route is tuned for faster direct answers while preserving multimodal and structured output support.","architecture":{"context_window":262144,"max_output_tokens":131072,"modality":"text+image"},"pricing":{"prompt":0.00013,"completion":0.0004,"cache_read":0.000065},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-3-4b-it","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 3 4B IT","description":"Gemma 3 has a large, 128K context window, multilingual support in over 140 languages, and is available in more sizes than previous versions.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0002006,"completion":0.0002006,"cache_read":0.0001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-3-27b-it","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 3 27B IT","description":"Gemma 3 has a large, 128K context window, multilingual support in over 140 languages, and is available in more sizes than previous versions.","architecture":{"context_window":128000,"max_output_tokens":96000,"modality":"text"},"pricing":{"prompt":0.0002992,"completion":0.0002992,"cache_read":0.0001496},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"gemma-3-12b-it","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Gemma 3 12B IT","description":"Gemma 3 has a large, 128K context window, multilingual support in over 140 languages, and is available in more sizes than previous versions.","architecture":{"context_window":128000,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.000272,"completion":0.000272,"cache_read":0.000136},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"eva-qwen2.5-72b-v0.2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"EVA-Qwen2.5-72B-v0.2","description":"A RP/storywriting specialist model, full-parameter finetune of Qwen2.5-72B on mixture of synthetic and natural data. It uses Celeste 70B 0.1 data mixture, greatly expanding it to improve versatility, creativity and flavor of the resulting model.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000799,"completion":0.000799,"cache_read":0.0003995},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"eva-qwen2.5-32b-v0.2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"EVA-Qwen2.5-32B-v0.2","description":"A RP/storywriting specialist model, full-parameter finetune of Qwen2.5-32B on mixture of synthetic and natural data. It uses Celeste 70B 0.1 data mixture, greatly expanding it to improve versatility, creativity and flavor of the resulting model.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000799,"completion":0.000799,"cache_read":0.0003995},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"eva-llama-3.33-70b-v0.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"EVA-LLaMA-3.33-70B-v0.1","description":"A RP/storywriting specialist model, full-parameter finetune of Llama-3.3-70B-Instruct on mixture of synthetic and natural data. It uses Celeste 70B 0.1 data mixture, greatly expanding it to improve versatility, creativity and flavor of the resulting model.","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.002006,"completion":0.002006,"cache_read":0.001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"eva-llama-3.33-70b-v0.0","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"EVA Llama 3.33 70B","description":"A RP/storywriting specialist model, full-parameter finetune of Llama-3.3-70B-Instruct on mixture of synthetic and natural data. It uses Celeste 70B 0.1 data mixture, greatly expanding it to improve versatility, creativity and flavor of the resulting model.","architecture":{"context_window":16384,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.002006,"completion":0.002006,"cache_read":0.001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"dracarys-72b-instruct","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Llama 3.1 70B Dracarys 2","description":"Llama 3.1 70b finetune that offers improvements on coding.","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.000493,"completion":0.000493,"cache_read":0.0002465},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"doubao-seed-character","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Doubao Seed Character","description":"ByteDance's character-focused Doubao Seed model for roleplay, persona consistency, dialogue, and creative character interactions. It supports text and image input with a 128k context window. Requests route through ZenMux to ByteDance; ZenMux does not publish a model-API zero-retention or training guarantee, so avoid sensitive data.","architecture":{"context_window":128000,"max_output_tokens":32768,"modality":"text+image"},"pricing":{"prompt":0.0001179,"completion":0.0002947,"cache_read":0.0000236},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"devstral-small-2505","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Mistral Devstral Small 2505","description":"OpenHands+Devstral is 100% local 100% open, and is SOTA for the category on SWE-Bench Verified: 46.8% accuracy.","architecture":{"context_window":32768,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.00006,"completion":0.00006,"cache_read":0.00003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"devstral-2-123b-instruct-2512","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Devstral 2 123B","description":"Devstral 2 123B is a 123 billion parameter model from Mistral AI optimized for coding and development tasks. Features advanced reasoning capabilities for software engineering workflows.","architecture":{"context_window":262144,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.0014,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v4-pro-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V4 Pro (Thinking)","description":"DeepSeek V4 Pro Thinking enables DeepSeek's chain-of-thought mode on the large-scale Mixture-of-Experts model with a 1M-token context window, built for advanced reasoning, coding, long-horizon agent workflows, knowledge, math, and software engineering tasks.","architecture":{"context_window":1048576,"max_output_tokens":384000,"modality":"text"},"pricing":{"prompt":0.0011,"completion":0.0022,"cache_read":0.00011},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v4-pro","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V4 Pro","description":"DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with a 1M-token context window, built for advanced reasoning, coding, long-horizon agent workflows, knowledge, math, and software engineering tasks.","architecture":{"context_window":1048576,"max_output_tokens":384000,"modality":"text"},"pricing":{"prompt":0.0011,"completion":0.0022,"cache_read":0.00011},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v4-flash-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V4 Flash (Thinking)","description":"DeepSeek V4 Flash Thinking enables DeepSeek's reasoning mode on the efficiency-optimized Mixture-of-Experts model with a 1M-token context window, built for fast inference, high-throughput workloads, reasoning, coding, and agent workflows.","architecture":{"context_window":1048576,"max_output_tokens":384000,"modality":"text"},"pricing":{"prompt":0.00014,"completion":0.00028,"cache_read":0.000028},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v4-flash","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V4 Flash","description":"DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with a 1M-token context window, built for fast inference, high-throughput workloads, reasoning, coding, and agent workflows.","architecture":{"context_window":1048576,"max_output_tokens":384000,"modality":"text"},"pricing":{"prompt":0.00014,"completion":0.00028,"cache_read":0.000028},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3.2-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3.2 Thinking","description":"DeepSeek V3.2 (thinking/reasoner mode) — official successor to V3.2-Exp. Reasoning-first model built for agents with GPT-5 level performance. Balanced inference vs. output length. First DeepSeek model with thinking-in-tool-use capability. FP8.","architecture":{"context_window":163000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.00028,"completion":0.00042,"cache_read":0.00014},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3.2-exp-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3.2 Exp Thinking","description":"Deepseek V3.2 Exp Thinking, Deepseek's latest model offering far better performance especially on longer contexts than its predecessors. Current flagship model by Deepseek. FP8.","architecture":{"context_window":163840,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.00028,"completion":0.00042,"cache_read":0.00014},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3.2-exp","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3.2 Exp","description":"Deepseek V3.2 Exp, Deepseek's latest model offering far better performance especially on longer contexts than its predecessors. Current flagship model by Deepseek. FP8.","architecture":{"context_window":163840,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.00028,"completion":0.00042,"cache_read":0.00014},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3.2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3.2","description":"DeepSeek V3.2 (non-thinking mode) — official successor to V3.2-Exp. Reasoning-first model built for agents with GPT-5 level performance. Balanced inference vs. output length for everyday use. First DeepSeek model with thinking-in-tool-use capability. FP8.","architecture":{"context_window":163000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.00028,"completion":0.00042,"cache_read":0.00014},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3.1-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3.1 Thinking","description":"Thinking enabled version of Deepseek V3.1. DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. It does better at tool calling and agent tasks, and has higher thinking efficiency than its predecessor. Quantized at FP8.","architecture":{"context_window":128000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0007,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3.1-terminus-thinking","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3.1 Terminus (Thinking)","description":"Thinking-enabled DeepSeek-V3.1-Terminus with improved language consistency, upgraded Code/Search Agents, and stronger stability and reliability versus V3.1. FP8.","architecture":{"context_window":128000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.00025,"completion":0.0007,"cache_read":0.000125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3.1-terminus","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3.1 Terminus","description":"DeepSeek-V3.1-Terminus. The latest update builds on V3.1's strengths while addressing key user feedback. Language consistency improvements (fewer CN/EN mix-ups, no random chars), stronger Code Agent \u0026 Search Agent performance, and more stable, reliable outputs across benchmarks. FP8.","architecture":{"context_window":128000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.00025,"completion":0.0007,"cache_read":0.000125},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3.1","description":"DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. It does better at tool calling and agent tasks, and has higher thinking efficiency than its predecessor. This is the non-thinking version. Quantized at FP8.","architecture":{"context_window":128000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.0007,"cache_read":0.0001},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-v3-0324","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek Chat 0324","description":"DeepSeek V3 0324, DeepSeek's 03 March 2025 V3 model, optimized for general-purpose tasks. Quantized at FP8.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0002,"completion":0.00077,"cache_read":0.000135},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-reasoner","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek Reasoner","description":"DeepSeek-R1 is now live and open source, rivaling OpenAI's Model o1.","architecture":{"context_window":64000,"max_output_tokens":65536,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.0017,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-r1-distill-qwen-32b-abliterated","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek R1 Qwen Abliterated","description":"Uncensored version of the Deepseek R1 Qwen 32B model","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0014,"completion":0.0014,"cache_read":0.0007},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-r1-distill-llama-70b-abliterated","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek R1 Llama 70B Abliterated","description":"Uncensored version of the Deepseek R1 Llama 70B model","architecture":{"context_window":16384,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0007,"completion":0.0007,"cache_read":0.00035},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-r1-0528","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek R1 0528","description":"The new (May 28th) Deepseek R1 model.","architecture":{"context_window":128000,"max_output_tokens":163840,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.0017,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-r1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek R1","description":"DeepSeek's R1 is a thinking model, scoring very well on all benchmarks at low cost. This version runs on open-source providers and never sends requests to DeepSeek directly.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0004,"completion":0.0017,"cache_read":0.0002},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-latest","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek Latest","description":"Compatibility alias that routes to the newest thinking DeepSeek model. Currently routes to DeepSeek V4 Pro Thinking.","architecture":{"context_window":1048576,"max_output_tokens":384000,"modality":"text"},"pricing":{"prompt":0.0011,"completion":0.0022,"cache_read":0.00011},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"deepseek-chat","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"DeepSeek V3/Deepseek Chat","description":"DeepSeek original V3 model, trained on nearly 15 trillion tokens, matches leading closed-source models at a far lower price. Quantized at FP8.","architecture":{"context_window":128000,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0001,"completion":0.000425,"cache_read":0.00005},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"cydonia-24b-v4.3","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"The Drummer Cydonia 24B v4.3","description":"Cydonia 24B v4.3 continues TheDrummer's Cydonia series with updated tuning on Mistral Small.","architecture":{"context_window":32768,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0001003,"completion":0.0001207,"cache_read":0.00005015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"cydonia-24b-v4.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"The Drummer Cydonia 24B v4.1","description":"Cydonia 24B v4.1 is the newest release of TheDrummer's Cydonia series, featuring improved performance and refined capabilities.","architecture":{"context_window":131072,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.00035,"completion":0.00055,"cache_read":0.00016},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"cydonia-24b-v4","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"The Drummer Cydonia 24B v4","description":"Cydonia 24B v4 is the latest iteration of TheDrummer's Cydonia series, a finetune of Mistral Small.","architecture":{"context_window":16384,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0002006,"completion":0.0002414,"cache_read":0.0001003},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"cydonia-24b-v2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"The Drummer Cydonia 24B v2","description":"Cydonia 24B v2 is a finetune of Mistral's latest 'Small' model (2501). Aliases: Cydonia 24B, Cydonia v2, Cydonia on that broken base.","architecture":{"context_window":16384,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0001003,"completion":0.0001207,"cache_read":0.00005015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"cogito-v1-preview-qwen-32b","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Cogito v1 Preview Qwen 32B","description":"32-B parameter reasoning model from DeepCogito (Qwen backbone) – strong general reasoning \u0026 coding at low price.","architecture":{"context_window":128000,"max_output_tokens":32768,"modality":"text"},"pricing":{"prompt":0.0018,"completion":0.0018,"cache_read":0.0009},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"claw-medium","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Claw Medium","description":"Currently points to MiniMax M2.7. Middle tier for OpenClaw-style agent work: medium intelligence and cost for capable everyday tool use.","architecture":{"context_window":204800,"max_output_tokens":131072,"modality":"text"},"pricing":{"prompt":0.000315,"completion":0.00126,"cache_read":0.0001575},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"anubis-70b-v1.1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Anubis 70B v1.1","description":"L3.3 finetune for roleplaying – updated v1.1 with improved reasoning.","architecture":{"context_window":131072,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00031,"completion":0.00031,"cache_read":0.000155},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"anubis-70b-v1","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Anubis 70B v1","description":"L3.3 finetune for roleplaying.","architecture":{"context_window":65536,"max_output_tokens":16384,"modality":"text"},"pricing":{"prompt":0.00031,"completion":0.00031,"cache_read":0.000155},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]},{"id":"amoral-gemma3-27b-v2","object":"model","created":1785441433,"owned_by":"helixmind","display_name":"Amoral Gemma3 27B v2","description":"Amoral Gemma3 27B v2 is a 27B parameter model that is a more advanced version of Gemma3 27B.","architecture":{"context_window":32768,"max_output_tokens":8192,"modality":"text"},"pricing":{"prompt":0.0003,"completion":0.0003,"cache_read":0.00015},"supported_endpoints":["/v1/chat/completions","/v1/messages","/v1/responses"]}]}
