News
AI Provider Index news.
Updates, product notes and practical coverage for AI platforms and operators.
Open Baseten Adds Hosted Inkling Small Model API Access
01 Aug 2026 · Ronald
Baseten Adds Hosted Inkling Small Model API Access
Baseten has added Inkling Small to its hosted Model APIs and dedicated-deployment service. The open-weight mixture-of-experts model has 276 billion total parameters, 12 billion active parameters, a one-million-token context window and native text, image and audio inputs.
Open Baseten Launches Distribution Platform for Closed Model Labs
31 Jul 2026 · Ronald
Baseten Launches Distribution Platform for Closed Model Labs
Baseten has launched a distribution and monetisation platform for developers of closed-weight AI models. Baseten for Model Labs combines managed inference, Model Library distribution and commercial support, giving specialist labs another route to production customers without building a complete serving business themselves.
Open Baseten Cuts Wan 2.2 Video Inference Below Three Seconds
17 Jul 2026 · Ronald
Baseten Cuts Wan 2.2 Video Inference Below Three Seconds
Baseten says its optimised Wan 2.2 runtime can generate a video clip in 2.75 seconds, 53.6 times faster than its baseline. A guarded public demo runs until 31 July, while the company acknowledges that four-step distillation can trade some visual quality for speed.
Open Baseten Adds NVIDIA Nemotron 3 Embed Models
16 Jul 2026 · Ronald
Baseten Adds NVIDIA Nemotron 3 Embed Models
Baseten has made NVIDIA's Nemotron 3 Embed 8B and 1B models available for dedicated inference. The pair targets enterprise and code retrieval with different balances of accuracy, throughput and indexing cost, giving developers a managed deployment option while published performance figures remain vendor-reported.
Open Baseten Adds Day-One Hosting for Inkling Multimodal Model
15 Jul 2026 · Ronald
Baseten Adds Day-One Hosting for Inkling Multimodal Model
Baseten has added day-one access to Thinking Machines Lab’s Inkling model through its managed Model APIs and Dedicated Inference service. The launch gives developers a hosted route to a very large open-weight, multimodal model, while leaving performance, cost and production suitability to be tested in real workloads.