LLM Model Support List¶
The following summarizes the list of HuggingFace models supported by the current version of the LLM converter tool and their corresponding running commands (LLM converter will continue to expand the adaptation range).
Note
The following models all require the use of officially released pre-trained or quantized weights.
--all_fp16 parameter is used for full FP16 conversion of non-AWQ/GPTQ models. It bypasses the quantization library and can directly generate offline models, but with lower inference speed and efficiency.
It is recommended to prioritize quantized weights such as AWQ/GPTQ to ensure the performance and efficiency of the converted model.
โ ๏ธ [Version Dependency Specification]
| Component | Required Version |
|---|---|
| LLM_Converter | Based on transformers 4.57.3 |
| Input Model | Must be an HF model trained with transformer 4.57.3 |
๐ Supported Model List¶
Update Time: 2025-01-09
Total Models: 34
๐ Model Classification¶
DeepSeek_R1¶
โ 1.5B
Practice
- Conversion command reference:
python3 convert_hf_to_sim.py \
-d /path/to/model/ \
-o /path/to/output_model/ \
--inputs /path/to/input.json \
--soc_version CHIP
-
For input.json construction method, please refer to the Constructing Quantization Calibration Data section in "SGS Quantized Model Conversion" chapter of Quick Start.
-
Running command reference:
python3 run_llm.py \
-d /path/to/output_model/ \
--prompt "Give me a short introduction to large language model." \
--soc_version CHIP
DeepSeek Janus Pro¶
โ 1B
โ ๏ธ Currently only Image-to-Text quantization is supported. Text-to-Image is not yet supported.
Practice
- Conversion command reference:
python3 convert_hf_to_sim.py \
-d /path/to/model/ \
-o /path/to/output_model/ \
--inputs /path/to/input.json \
--visual_input_formats YUV_NV12 \
--soc_version CHIP
-
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command reference:
python3 run_llm.py \
-d /path/to/output_model/ \
-t /path/to/model/ \
--image /path/to/jpg/ \
--no_streamer \
--prompt "Please describe this image" \
--soc_version CHIP
Qwen2.5¶
โ 0.5B
โ 1.5B
โ 1.5B-Int4
โ 1.5B-Int4
โ 1.5B-Int8
โ 3B
Practice
-
Conversion command reference (AWQ/GPTQ models do not require
--inputsparameter):python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --soc_version CHIP -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model." \ --soc_version CHIP
Qwen3¶
โ 0.6B
- ๐ Link: https://huggingface.co/Qwen/Qwen3-0.6B
โ 1.7B
- ๐ Link: https://huggingface.co/Qwen/Qwen3-1.7B
โ 4B
- ๐ Link: https://huggingface.co/Qwen/Qwen3-4B
Practice
- Conversion command reference:
python3 convert_hf_to_sim.py \
-d /path/to/model/ \
-o /path/to/output_model/ \
--inputs /path/to/input.json \
--soc_version CHIP
-
For input.json construction method, please refer to the Constructing Quantization Calibration Data section in "SGS Quantized Model Conversion" chapter of Quick Start.
-
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model." \ --soc_version CHIP
Qwen2.5-VL¶
โ 3B
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --imgsz 672 392 --n_tokens 512 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --image /path/to/jpg/ \ --processor_decode --no_streamer \ --prompt "Please describe this image" \ --soc_version CHIP
Qwen2.5-Omni¶
โ 3B
- ๐ Link: https://huggingface.co/Qwen/Qwen2.5-Omni-3B
โ ๏ธ SGS quantization is not yet supported.
Practice
-
Conversion command reference:
python3 ./convert_hf_to_sim.py \ -d /path/to/model/ --n_tokens 64 \ --imgsz 224 224 \ --videosz 8 3 308 308 \ --soc_version CHIP -
Running command reference:
python3 ./run_llm.py \ -d /path/to/output_model/ \ --prompt "Describe the input" \ --image /path/to/image/ \ --video /path/to/video/ \ --wav_file /path/to/wav/ \ --use_audio_in_video \ --soc_version CHIP
Qwen3-VL¶
โ 2B
Practice
-
Conversion command reference:
python3 ../Tool/Scripts/LLM_Converter/convert_hf_to_sim.py \ -d /path/to/Qwen/Qwen3-VL-2B-Instruct/ \ -o output_models \ --visual_input_formats YUV_NV12 \ --imgsz 384 384 \ --videosz 8 3 384 384 \ --num_soc 2 \ --inputs inputs.json -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command reference:
python3 ../Tool/Scripts/LLM_Converter/run_llm.py \ -t /path/to/Qwen/Qwen3-VL-2B-Instruct/ \ -d output_models \ --prompt "Briefly describe this image" \ --image /path/to/image.jpg \ --soc_version CHIP
opus-mt-zh-en¶
โ 140M
Practice
-
Conversion command:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs input.json \ -q Q15 Q10 Q25 \ --max_length 512 \ --encoder_tokens 128 \ --soc_version CHIP -
input.json format example:
[ {"prompt": "ไธญๅฝๅๅคงๅ่ๆๅชไบ๏ผ"}, {"prompt": "่ฏท็ๆไธๆฎต็งๆๆน้ข็ๆๆฌ๏ผ512ไธชๅญใ"}, {"prompt": "5+8*10็็ปๆๆฏ๏ผ"}, {"prompt": "่ฏทๅฐๅฆไธไธญๆ็ฟป่ฏไธบ่ฑๆ๏ผๆๆ่ช่ฟๆนๆฅ๏ผไธไบฆไนไนใ"} ] -
Running command:
python3 run_llm.py \ -d /path/to/output_model/ \ --no_streamer \ --prompt "ไฝ ๅฅฝ๏ผๆๆฏไธญๅฝไบบ๏ผ" \ --soc_version CHIP
InternVL2.5-MP0¶
โ 1B
Practice
-
Conversion command:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command:
python3 run_llm.py \ -d /path/to/output_model/ \ --image /path/to/jpg/ \ --prompt "Please describe this image" \ --soc_version CHIP
โ 2B
Practice
-
Conversion command:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command:
python3 run_llm.py \ -d /path/to/output_model/ \ --image /path/to/jpg/ \ --prompt "Please describe this image" \ --soc_version CHIP
InternVL2.5¶
โ 1B
โ 2B
Practice
-
Conversion command:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command:
python3 run_llm.py \ -d /path/to/output_model/ \ --image /path/to/jpg/ \ --prompt "Please describe this image" \ --soc_version CHIP
Gemma¶
โ 2B
- ๐ Link: https://huggingface.co/google/gemma-2b
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --soc_version CHIP -
For input.json construction method, please refer to the Constructing Quantization Calibration Data section in "SGS Quantized Model Conversion" chapter of Quick Start.
-
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model." \ --soc_version CHIP
Gemma2¶
โ 2B
- ๐ Link: https://huggingface.co/google/gemma-2-2b-it
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --soc_version CHIP -
For input.json construction method, please refer to the Constructing Quantization Calibration Data section in "SGS Quantized Model Conversion" chapter of Quick Start.
-
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model." \ --soc_version CHIP
Gemma3¶
โ 1B
- ๐ Link: https://huggingface.co/google/gemma-3-1b-it
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --soc_version CHIP -
For input.json construction method, please refer to the Constructing Quantization Calibration Data section in "SGS Quantized Model Conversion" chapter of Quick Start.
-
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model." \ --soc_version CHIP
Llama 3.2¶
โ 1B
โ 3B
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --soc_version CHIP -
For input.json construction method, please refer to the Constructing Quantization Calibration Data section in "SGS Quantized Model Conversion" chapter of Quick Start.
-
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model." \ --soc_version CHIP
Phi 3 Mini¶
โ 3.8B
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --soc_version CHIP -
For input.json construction method, please refer to the Constructing Quantization Calibration Data section in "SGS Quantized Model Conversion" chapter of Quick Start.
-
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model." \ --soc_version CHIP
Phi 4 Mini¶
โ 3.8B
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --soc_version CHIP -
For input.json construction method, please refer to the Constructing Quantization Calibration Data section in "SGS Quantized Model Conversion" chapter of Quick Start.
-
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model." \ --soc_version CHIP
MiniMind2¶
โ 104M
- ๐ Link: https://huggingface.co/jingyaogong/MiniMind2
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --soc_version CHIP -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --prompt "Give me a short introduction to large language model."
MiniMind2-V¶
โ 104M
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model/ \ -o /path/to/output_model/ \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model/ \ --image /path/to/demo.jpg \ --prompt "Describe this image shortly." \ --soc_version CHIP
Florence-2-base¶
โ 0.2B
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model \ -o /path/to/output_model \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ -q Q15 Q10 Q25 Q25 \ --soc_version CHIP -
input.json format example:
[ {"prompt": "<OD>", "image": "/path/to/image.jpg"}, {"prompt": "<CAPTION>", "image": "/path/to/image.jpg"} ] -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model \ --prompt "<OD>" \ --image /path/to/image \ --no_streamer \ --processor_postprocess_generate \ --processor_decode \ --soc_version CHIP
SmolVLM-256M¶
โ 256M
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model \ -o /path/to/output_model \ --n_tokens 2048 \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model \ -t /path/to/model \ --image /path/to/demo.jpg \ --prompt "Describe this image shortly." \ --no_streamer \ --soc_version CHIP
SmolVLM2-500M¶
โ 500M
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model \ -o /path/to/output_model \ --n_tokens 2048 \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model \ -t /path/to/model \ --image /path/to/demo.jpg \ --prompt "Describe this image shortly." \ --no_streamer \ --soc_version CHIP
FastVLM¶
โ 0.5B
- ๐ Link: https://huggingface.co/apple/FastVLM-0.5B
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model \ -o /path/to/output_model \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --max_length 512 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model \ -t /path/to/model \ --image /path/to/demo.jpg \ --prompt "Describe this image shortly." \ --no_streamer \ --soc_version CHIP
โ 1.5B
- ๐ Link: https://huggingface.co/apple/FastVLM-1.5B
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model \ -o /path/to/output_model \ --inputs /path/to/input.json \ --visual_input_formats YUV_NV12 \ --max_length 512 \ --soc_version CHIP -
input.json format example:
[ { "prompt": "Briefly describe this image", "image": "/path/to/image.jpg" } ] -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model \ -t /path/to/model \ --image /path/to/demo.jpg \ --prompt "Describe this image shortly." \ --no_streamer \ --soc_version CHIP
whisper-tiny¶
โ 37.8M
- ๐ Link: https://huggingface.co/openai/whisper-tiny
Practice
-
Conversion command reference:
python3 convert_hf_to_sim.py \ -d /path/to/model \ -o /path/to/output_model \ --inputs /path/to/input.json \ -q Q15 Q10 Q25 \ --max_length 512 \ --soc_version CHIP -
input.json format example:
[ {"prompt": null, "wav_file": ["/path/to/new.wav"]}, {"prompt": null, "wav_file": ["/path/to/take_photo.wav"]} ] -
Running command reference:
python3 run_llm.py \ -d /path/to/output_model \ --wav_file /path/to/wav \ --no_streamer \ --host 10.44.16.18 \ --port 3333 \ --model_onboard_dir /path/to/board \ --soc_version CHIP