数据集 / unsloth/Qwen3-30B-A3B-128K-GGUF

unsloth/Qwen3-30B-A3B-128K-GGUF 已完整同步

See our collection for all versions of Qwen3 including GGUF, 4-bit & 16-bit formats.

Learn to run Qwen3 correctly - Read our Guide.

Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants.

Run & Fine-tune Qwen3 with Unsloth!

New updated quants

Unsloth supports Free Notebooks Performance Memory use
Qwen3 (14B) ▶️ Start on Colab 3x faster 70% less
GRPO with Qwen3 (8B) ▶️ Start on Colab 3x faster 80% less
Llama-3.2 (3B) ▶️ Start on Colab 2.4x faster 58% less
Llama-3.2 (11B vision) ▶️ Start on Colab 2x faster 60% less
Qwen2.5 (7B) ▶️ Start on Colab 2x faster 60% less
Phi-4 (14B) ▶️ Start on Colab 2x faster 50% less

To Switch Between Thinking and Non-Thinking

If you are using llama.cpp, Ollama, Open WebUI etc., you can add /think and /no_think to user prompts or system messages to switch the model's thinking mode from turn to turn. The model will follow the most recent instruction in multi-turn conversations.

Here is an example of multi-turn conversation:

> Who are you /no_think

<think>

</think>

I am Qwen, a large-scale language model developed by Alibaba Cloud. [...]

> How many 'r's are in 'strawberries'? /think

<think>
Okay, let's see. The user is asking how many times the letter 'r' appears in the word "strawberries". [...]
</think>

The word strawberries contains 3 instan

33 个文件

浏览文件