vllm.model_executor.models ¶
Modules:
-
AXK1–Inference-only A.X K1 model.
-
adapters– -
afmoe–Inference-only AfMoE model compatible with HuggingFace weights.
-
apertus–Inference-only Apertus model compatible with HuggingFace weights.
-
arcee– -
arctic–Inference-only Snowflake Arctic model.
-
aria– -
audioflamingo3– -
bagel–Inference-only BAGEL model compatible with HuggingFace weights.
-
bailing_moe–Inference-only BailingMoE model compatible with HuggingFace weights.
-
bailing_moe_linear– -
bailing_moe_mtp–Inference-only Bailing MoE v2.5 MTP model.
-
bailing_moe_v3–vLLM implementation for BailingMoeV3ForCausalLM.
-
bailing_moe_v3_mtp–Inference-only Bailing MoE v3 MTP model.
-
bee– -
bert– -
blip–Minimal implementation of BlipVisionModel intended to be only used
-
blip2– -
bloom–Inference-only BLOOM model compatible with HuggingFace weights.
-
chameleon– -
chatglm–Inference-only ChatGLM model compatible with THUDM weights.
-
cheers–Inference-only Cheers (UMM) model compatible with HuggingFace weights.
-
clip– -
cohere2_moe– -
cohere2_vision–Command-A-Vision (Cohere2Vision) multimodal model implementation for vLLM.
-
cohere_asr– -
cohere_eagle– -
colbert–ColBERT late interaction model for retrieval and reranking.
-
colmodernvbert–ColModernVBERT: multimodal late-interaction retrieval model.
-
colpali–ColPali late interaction model for multi-modal retrieval and reranking.
-
colqwen3–ColQwen3 late interaction model for multi-modal retrieval and reranking.
-
colqwen3_5–ColQwen3.5 late interaction model for multi-modal retrieval and reranking.
-
commandr–PyTorch Cohere model.
-
config– -
conformer_encoder–Shared Conformer encoder components for FireRedASR2 and FireRedLID.
-
cosmos3_edge– -
dbrx– -
deepencoder– -
deepencoder2– -
deepseek_eagle3–Eagle3 speculative decoding model for DeepseekV2/V3 with MLP (no MoE).
-
deepseek_mtp– -
deepseek_ocr–Inference-only Deepseek-OCR model compatible with HuggingFace weights.
-
deepseek_ocr2–Inference-only Deepseek-OCR model compatible with HuggingFace weights.
-
deepseek_v2–Inference-only DeepseekV2/DeepseekV3 model.
-
deepseek_vl2–Inference-only Deepseek-VL2 model compatible with HuggingFace weights.
-
diffusion_gemma–DiffusionGemma model, ModelState, and Sampler for vLLM.
-
dots_ocr– -
eagle2_5_vl– -
ernie45–Inference-only Erine model compatible with HuggingFace weights.
-
ernie45_moe–Inference-only ErineMoE model compatible with HuggingFace weights.
-
ernie45_vl–Inference-only Ernie VL model compatible with HuggingFace weights.
-
ernie45_vl_moe–Inference-only Erine VL model compatible with HuggingFace weights.
-
ernie_mtp–Inference-only Ernie-MTP model.
-
exaone–Inference-only Exaone model compatible with HuggingFace weights.
-
exaone4–Inference-only Exaone model compatible with HuggingFace weights.
-
exaone4_5–Inference-only EXAONE-4.5 model compatible with HuggingFace weights.
-
exaone4_5_mtp–Inference-only EXAONE-4_5 MTP model.
-
exaone_moe–Inference-only K-EXAONE-236B-A22B model compatible with HuggingFace weights.
-
exaone_moe_mtp–Inference-only ExaoneMoe MTP model.
-
extract_hidden_states–Hidden States Extractor Model.
-
fairseq2_llama–Llama model for fairseq2 weights.
-
falcon–PyTorch Falcon model.
-
falcon_h1–Inference-only FalconH1 model.
-
fireredasr2– -
fireredlid–FireRedLID – Language Identification model adapted for vLLM.
-
flex_olmo–Inference-only FlexOlmo model compatible with HuggingFace weights.
-
funasr– -
funaudiochat–Inference-only FunAudioChat model compatible with HuggingFace weights.
-
gemma–Inference-only Gemma model compatible with HuggingFace weights.
-
gemma3_mm– -
gemma3n– -
gemma3n_audio_utils–Lightweight utility functions for Gemma3n audio processing.
-
gemma3n_mm– -
gemma4–Gemma 4 model implementation for vLLM.
-
gemma4_dspark–Gemma4 DSpark draft model for speculative decoding.
-
gemma4_mm–Gemma 4 multimodal model (image + audio + video support).
-
gemma4_mtp–Inference-only Gemma4 MTP (Multi-Token Prediction) model.
-
gemma4_unified–Gemma 4 Unified multimodal model (encoder-free image + audio + video).
-
glm–Inference-only HF format GLM-4 model compatible with THUDM weights.
-
glm4–Inference-only GLM-4-0414 model compatible with HuggingFace weights.
-
glm4_1v–Inference-only GLM-4.1V & GLM-4.6V-Flash, AutoGLM-Phone-9B model
-
glm4_moe–Inference-only GLM-4.5, GLM-4.6, GLM-4.7 model
-
glm4_moe_lite–Inference-only GLM-4.7-Flash model compatible with HuggingFace weights.
-
glm4_moe_lite_mtp–Inference-only GLM-4.7-Flash MTP model compatible with HuggingFace weights.
-
glm4_moe_mtp–Inference-only GLM-4.5, GLM-4.6, GLM-4.7 MTP
-
glm4v–Inference-only CogAgent model compatible with THUDM weights.
-
glm_ocr–Inference-only GLM-OCR model compatible with HuggingFace weights.
-
glm_ocr_mtp–Inference-only GLM-OCR MTP model compatible with HuggingFace weights.
-
glmasr– -
glmasr_utils– -
gpt2–Inference-only GPT-2 model compatible with HuggingFace weights.
-
gpt_j–Inference-only GPT-J model compatible with HuggingFace weights.
-
gpt_neox–Inference-only GPT-NeoX model compatible with HuggingFace weights.
-
granite–Inference-only IBM Granite model compatible with HuggingFace weights.
-
granite4_vision–vLLM implementation of Granite 4 Vision.
-
granite_speech–Inference-only IBM Granite speech model.
-
granite_speech_plus–Inference-only IBM Granite Speech Plus model.
-
granitemoe–Inference-only GraniteMoe model.
-
granitemoehybrid–Inference-only GraniteMoeHybrid model.
-
granitemoeshared–Inference-only GraniteMoeShared model.
-
gritlm– -
hrm_text–HRM-Text: Hierarchical Reasoning Model — Text variant.
-
hunyuan_v1–Inference-only HunYuan model compatible with HuggingFace weights.
-
hunyuan_vision–Inference-only HunYuan-VL model compatible with HuggingFace weights.
-
hy_v3–Inference-only HY model compatible with HuggingFace weights.
-
hy_v3_mtp–Inference-only HY V3 MTP model compatible with HuggingFace weights.
-
hyperclovax–Inference-only HyperCLOVAX model compatible with HuggingFace weights.
-
hyperclovax_vision– -
hyperclovax_vision_v2–HyperCLOVAX V2 (32B Think Model) Implementation.
-
idefics2_vision_model–PyTorch Idefics2 model.
-
idefics3–Inference-only Idefics3 model compatible with HuggingFace weights.
-
interfaces– -
interfaces_base– -
intern_vit– -
interns1– -
interns1_pro–Inference-only InternS1Pro model compatible with HuggingFace weights.
-
interns1_vit– -
interns2_mobius–Inference-only Intern-S2-Mobius model.
-
internvl– -
iquest_loopcoder–Inference-only LoopCoder model compatible with HuggingFace weights.
-
isaac– -
jais2–Inference-only Jais2 model compatible with HuggingFace weights.
-
jamba–Inference-only Jamba model.
-
jina– -
kanana_v– -
keye– -
keye_vl1_5– -
kimi_audio–Inference-only Kimi-Audio model compatible with HuggingFace weights.
-
kimi_k25–Kimi-K2.5 Model Implementation for vLLM.
-
kimi_k25_vit–Vision tower implementation for Kimi-K2.5 model.
-
kimi_vl– -
laguna–Inference-only Laguna model compatible with HuggingFace weights.
-
laguna_dflash–DFlash speculator for Laguna target models.
-
lfm2– -
lfm2_moe– -
lfm2_siglip2–Implementation of Siglip2VisionModel intended to be only used
-
lfm2_vl– -
llama–Inference-only LLaMA model compatible with HuggingFace weights.
-
llama4–Inference-only LLaMA model compatible with HuggingFace weights.
-
llava– -
llava_next– -
llava_next_video– -
llava_onevision– -
llava_onevision2–Inference-only LLaVA-OneVision-2 (OV2) model for vLLM.
-
longcat_flash–Inference-only Flash model compatible with HuggingFace weights.
-
longcat_flash_mtp– -
longcat_flash_ngram–Inference-only LongCat-Flash-Lite (n-gram embedding) model.
-
mamba–PyTorch MAMBA model.
-
mamba2–PyTorch MAMBA2 model.
-
medusa– -
mellum– -
midashenglm–Inference-only MiDashengLM model compatible with HuggingFace weights.
-
mimo–Inference-only MiMo model compatible with HuggingFace weights.
-
mimo_audio–MiMo audio: tokenizer, encoding utilities, and audio encoder.
-
mimo_mtp–Inference-only MiMo-MTP model.
-
mimo_v2– -
mimo_v2_mtp–Inference-only MiMo-V2 MTP (Multi-Token Prediction) draft model.
-
mimo_v2_omni– -
minicpm–Inference-only MiniCPM model compatible with HuggingFace weights.
-
minicpm3–Inference-only MiniCPM3 model compatible with HuggingFace weights.
-
minicpm_eagle–Inference-only EagleMiniCPM model compatible with HuggingFace weights.
-
minicpmo–Inference-only MiniCPM-O model compatible with HuggingFace weights.
-
minicpmv–Inference-only MiniCPM-V model compatible with HuggingFace weights.
-
minicpmv4_6–Inference-only MiniCPM-V 4.6 model (MiniCPMV4_6ForConditionalGeneration).
-
minimax_m2–Inference-only MiniMaxM2 model.
-
mistral–Mistral adaptation of the LLaMA architecture.
-
mistral3– -
mixtral–Inference-only Mixtral model.
-
mllama4– -
mlp_speculator– -
molmo– -
molmo2– -
moondream3–Inference-only Moondream3 model implementation.
-
moonvit– -
moss_audio–Inference-only MOSS-Audio model compatible with HuggingFace weights.
-
moss_transcribe_diarize–Inference-only MOSS-Transcribe-Diarize ASR model.
-
muse_glimmer–Inference-only MuseGlimmer multimodal model for vLLM.
-
nano_nemotron_vl– -
nemotron–Inference-only Nemotron model compatible with HuggingFace weights.
-
nemotron_h–Inference-only NemotronH model.
-
nemotron_h_mtp–NemotronH-MTP model with attention layers.
-
nemotron_nas–Inference-only deci model compatible with HuggingFace weights.
-
nemotron_parse– -
nemotron_vl– -
olmo3–Inference-only OLMo3 model compatible with HuggingFace weights.
-
olmo_hybrid–Inference-only OLMo Hybrid model compatible with HuggingFace weights.
-
olmoe–Inference-only OLMoE model compatible with HuggingFace weights.
-
openai_privacy_filter–Inference-only OpenAI Privacy Filter model.
-
opencua–Inference-only OpenCUA-7B model compatible with HuggingFace weights.
-
openpangu_mtp– -
openpangu_vl– -
openvla– -
opt–Inference-only OPT model compatible with HuggingFace weights.
-
orion–Inference-only Orion-14B model compatible with HuggingFace weights.
-
ovis–PyTorch Ovis model.
-
ovis2_5–PyTorch Ovis model.
-
paddleocr_vl– -
paligemma– -
parakeet–Modules below used for the audio encoder component in: models/nano_nemotron_vl.py
-
param2moe– -
phi–Inference-only Phi-1.5 model compatible with HuggingFace weights.
-
phi3–Inference-only Phi3 model code inherit from Llama.py
-
phi3v– -
phi4mm– -
phi4mm_audio– -
phi4mm_utils– -
phi4siglip–vLLM support for microsoft/Phi-4-reasoning-vision-15B.
-
phimoe–Inference-only PhiMoE model.
-
pixtral– -
plamo3–Inference-only PLaMo3 model.
-
qianfan_ocr– -
qwen2–Inference-only Qwen2 model compatible with HuggingFace weights.
-
qwen2_5_omni_thinker–Inference-only Qwen2.5-Omni model (thinker part).
-
qwen2_5_vl–Inference-only Qwen2.5-VL model compatible with HuggingFace weights.
-
qwen2_audio–Inference-only Qwen2-Audio model compatible with HuggingFace weights.
-
qwen2_moe–Inference-only Qwen2MoE model compatible with HuggingFace weights.
-
qwen2_rm–Inference-only Qwen2-RM model compatible with HuggingFace weights.
-
qwen2_vl–Inference-only Qwen2-VL model compatible with HuggingFace weights.
-
qwen3–Inference-only Qwen3 model compatible with HuggingFace weights.
-
qwen3_5–Inference-only Qwen3.5 Series compatible with HuggingFace weights.
-
qwen3_5_mtp–Inference-only Qwen3_5 MTP model.
-
qwen3_asr–Inference-only Qwen3-ASR model.
-
qwen3_asr_forced_aligner–Inference-only Qwen3-ASR ForcedAligner model (token classification).
-
qwen3_asr_realtime–Inference-only Qwen3-ASR realtime model.
-
qwen3_dflash– -
qwen3_dspark–Qwen3 DSpark draft model for semi-autoregressive drafting.
-
qwen3_moe–Inference-only Qwen3MoE model compatible with HuggingFace weights.
-
qwen3_next–Inference-only Qwen3Next model.
-
qwen3_next_mtp–Inference-only Qwen3Next MTP model.
-
qwen3_omni_moe_thinker–Inference-only Qwen3-Omni-Moe model (thinker part).
-
qwen3_vl–Inference-only Qwen3VL model compatible with HuggingFace weights.
-
qwen3_vl_moe–Inference-only Qwen3-VL-MoE model compatible with HuggingFace weights.
-
radio– -
registry–Whenever you add an architecture to this page, please also update
-
roberta– -
sarvam– -
seed_oss–Inference-only SeedOss model compatible with HuggingFace weights.
-
siglip– -
siglip2navit–Implementation of SiglipVisionModel intended to be only used
-
skyworkr1v– -
solar–Inference-only Solar model compatible with HuggingFace weights.
-
stablelm–Inference-only StableLM (https://github.com/Stability-AI/StableLM)
-
step1–Shared Step decoder blocks and the Step1 text model.
-
step3_text–Inference-only Jurassic model.
-
step3_vl– -
step3p5–Inference-only Jurassic model.
-
step3p5_mtp– -
step3p7–Inference-only Jurassic model.
-
step_vl–This is basically a copy from perception_models/core/vision_encoder/pe.py
-
terratorch–Wrapper around
Terratorchmodels -
transformers–Wrapper around
transformersmodels -
ultravox–PyTorch Ultravox model.
-
unlimited_ocr–Inference-only Unlimited-OCR model compatible with HuggingFace weights.
-
utils– -
vision– -
voxtral– -
voxtral_realtime– -
voyage– -
whisper– -
whisper_causal– -
zamba2–PyTorch Zamba2 model implementation for vLLM.
Classes:
-
HasInnerState–The interface required for all models that has inner state.
-
SupportsLoRA–The interface required for all models that support LoRA.
-
SupportsMRoPE–The interface required for all models that support M-RoPE.
-
SupportsMultiModal–The interface required for all multi-modal models.
-
SupportsMultiModalEmbeddings–The interface for models that can merge external multimodal embeddings.
-
SupportsPP–The interface required for all models that support pipeline parallel.
-
SupportsTranscription–The interface required for all models that support transcription.
-
VllmModelForPooling–The interface required for all pooling models in vLLM.
-
VllmModelForTextGeneration–The interface required for all generative models in vLLM.
HasInnerState ¶
Bases: Protocol
The interface required for all models that has inner state.
Attributes:
-
has_inner_state(Literal[True]) –A flag that indicates this model has inner state.
Source code in vllm/model_executor/models/interfaces.py
has_inner_state = True class-attribute ¶
A flag that indicates this model has inner state. Models that has inner state usually need access to the scheduler_config for max_num_seqs, etc. True for e.g. both Mamba and Jamba.
SupportsLoRA ¶
Bases: Protocol
The interface required for all models that support LoRA.
Attributes:
-
supports_lora(Literal[True]) –A flag that indicates this model supports LoRA.
Source code in vllm/model_executor/models/interfaces.py
supports_lora = True class-attribute ¶
A flag that indicates this model supports LoRA.
Note
There is no need to redefine this flag if this class is in the MRO of your model class.
SupportsMRoPE ¶
Bases: Protocol
The interface required for all models that support M-RoPE.
Methods:
-
get_mrope_input_positions–Get M-RoPE input positions and delta value for this specific model.
Attributes:
-
supports_mrope(Literal[True]) –A flag that indicates this model supports M-RoPE.
Source code in vllm/model_executor/models/interfaces.py
supports_mrope = True class-attribute ¶
A flag that indicates this model supports M-RoPE.
Note
There is no need to redefine this flag if this class is in the MRO of your model class.
get_mrope_input_positions(input_tokens, mm_features) ¶
Get M-RoPE input positions and delta value for this specific model.
This method should be implemented by each model that supports M-RoPE to provide model-specific logic for computing input positions.
Parameters:
-
(input_tokens¶list[int]) –List of input token IDs
-
(mm_features¶list[MultiModalFeatureSpec]) –Information about each multi-modal data item
Returns:
-
Tensor–Tuple of
(llm_positions, mrope_position_delta) -
int–- llm_positions: Tensor of shape
[3, num_tokens]with T/H/W positions
- llm_positions: Tensor of shape
-
tuple[Tensor, int]–- mrope_position_delta: Delta for position calculations
Source code in vllm/model_executor/models/interfaces.py
SupportsMultiModal ¶
Bases: SupportsMultiModalEmbeddings, Protocol
The interface required for all multi-modal models.
Methods:
-
configure_mm_token_handling–Check if any multimodal tokens are out of vocabulary. If so, we will
-
embed_input_ids–Apply token embeddings to
input_ids. -
embed_multimodal–Returns multimodal embeddings generated from multimodal kwargs
-
get_language_model–Returns the underlying language model used for text generation.
-
get_mm_lora_token_counts–Return
(tower_tokens, connector_tokens)for multimodal LoRA mappings. -
get_num_mm_connector_tokens–Implement this function to enable LoRA support
-
get_num_mm_encoder_tokens–Implement this function to enable LoRA support
-
get_placeholder_str–Get the placeholder text for the
ithmodalityitem in the prompt.
Attributes:
-
requires_raw_input_tokens(bool) –A flag that indicates this model processes input id tokens
-
supports_encoder_tp_data(bool) –A flag that indicates whether this model supports
-
supports_mm_device_do_normalize(bool) –A flag that indicates whether this model supports
-
supports_multimodal(Literal[True]) –A flag that indicates this model supports multi-modal inputs.
-
supports_multimodal_raw_input_only(bool) –A flag that indicates this model supports multi-modal inputs and processes
-
supports_tower_connector_lora(bool) –A flag that indicates whether this model supports
Source code in vllm/model_executor/models/interfaces.py
139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 | |
_has_oov_mm_tokens = False class-attribute instance-attribute ¶
In general, this should be set at init time by invoking configure_mm_token_handling models & passing all potentially OOV multimodal tokens.
_language_model_names = [] class-attribute instance-attribute ¶
Set internally by _mark_language_model.
_processor_factory class-attribute ¶
Set internally by MultiModalRegistry.register_processor.
_tower_model_names = [] class-attribute instance-attribute ¶
Set internally by _mark_tower_model.
requires_raw_input_tokens = False class-attribute ¶
A flag that indicates this model processes input id tokens in their raw form and not input embeddings.
supports_encoder_tp_data = False class-attribute ¶
A flag that indicates whether this model supports multimodal_config.mm_encoder_tp_mode="data".
supports_mm_device_do_normalize = False class-attribute ¶
A flag that indicates whether this model supports multimodal_config.mm_device_do_normalize.
supports_multimodal = True class-attribute ¶
A flag that indicates this model supports multi-modal inputs.
Note
There is no need to redefine this flag if this class is in the MRO of your model class.
supports_multimodal_raw_input_only = False class-attribute ¶
A flag that indicates this model supports multi-modal inputs and processes them in their raw form and not embeddings.
supports_tower_connector_lora = False class-attribute ¶
A flag that indicates whether this model supports lora_config.enable_tower_connector_lora.
_mark_composite_model(vllm_config, *, language_targets, tower_targets) ¶
Composite wrapper over _mark_language_model and _mark_tower_model by modality.
Source code in vllm/model_executor/models/interfaces.py
_mark_language_model(vllm_config, *, targets=None) ¶
Mark each child module that was assigned to this model during this context as a language model component.
Language model components are automatically skipped in --mm-encoder-only mode.
If targets is set, instead include descendants that are an instance of targets, even if they aren't direct children.
Source code in vllm/model_executor/models/interfaces.py
_mark_tower_model(vllm_config, modalities, *, targets=None) ¶
Mark each child module that was assigned to this model during this context as a tower model component.
Tower model components are automatically skipped when --limit-mm-per-prompt is set to zero for all of their modalities.
If targets is set, instead include descendants that are an instance of targets, even if they aren't direct children.
Source code in vllm/model_executor/models/interfaces.py
configure_mm_token_handling(vocab_size, mm_token_ids) ¶
Check if any multimodal tokens are out of vocabulary. If so, we will explicitly mask all multimodal tokens out when computing text embeddings, since the multimodal embeddings will be scattered over the results.
Source code in vllm/model_executor/models/interfaces.py
embed_input_ids(input_ids, multimodal_embeddings=None, *, is_multimodal=None) ¶
Apply token embeddings to input_ids.
If multimodal_embeddings is passed, scatter them into input_ids according to the mask is_multimodal.
NOTE: If this model has multimodal tokens that are of vocabulary (i.e., self._has_oov_mm_tokens=True), the input_ids will be copied and masked to 0 during the forward pass for the text embeddings.
Source code in vllm/model_executor/models/interfaces.py
embed_multimodal(**kwargs) ¶
Returns multimodal embeddings generated from multimodal kwargs to be merged with text embeddings.
Note
The returned multimodal embeddings must be in the same order as the appearances of their corresponding multimodal data item in the input prompt.
Source code in vllm/model_executor/models/interfaces.py
get_language_model() ¶
Returns the underlying language model used for text generation.
This is typically the torch.nn.Module instance responsible for processing the merged multimodal embeddings and producing hidden states
Returns:
-
VllmModel–torch.nn.Module: The core language model component.
Source code in vllm/model_executor/models/interfaces.py
get_mm_lora_token_counts(*, modality, mm_kwargs, num_mm_embeds) ¶
Return (tower_tokens, connector_tokens) for multimodal LoRA mappings.
MM LoRA uses these counts to build adapter mappings for the tower and connector forwards. Models with multiple modalities can override this when each modality has different encoder padding or pooling behavior.
Source code in vllm/model_executor/models/interfaces.py
get_num_mm_connector_tokens(num_vision_tokens) ¶
Implement this function to enable LoRA support for the connector module of the multi-modal model. Given the number of vision tokens, output the number of multi-modal connector tokens.
Source code in vllm/model_executor/models/interfaces.py
get_num_mm_encoder_tokens(num_image_tokens) ¶
Implement this function to enable LoRA support for the tower module of the multi-modal model. Given the number of image tokens, output the number of multi-modal encoder tokens.
Source code in vllm/model_executor/models/interfaces.py
get_placeholder_str(modality, i) classmethod ¶
Get the placeholder text for the ith modality item in the prompt.
SupportsMultiModalEmbeddings ¶
Bases: Protocol
The interface for models that can merge external multimodal embeddings.
Source code in vllm/model_executor/models/interfaces.py
SupportsPP ¶
Bases: Protocol
The interface required for all models that support pipeline parallel.
Methods:
-
forward–Accept
IntermediateTensorswhen
Attributes:
-
make_empty_intermediate_tensors(_MakeEmptyIntermediateTensors) –Called when PP rank > 0 for profiling purposes.
-
supports_pp(Literal[True]) –A flag that indicates this model supports pipeline parallel.
Source code in vllm/model_executor/models/interfaces.py
make_empty_intermediate_tensors instance-attribute ¶
Called when PP rank > 0 for profiling purposes.
supports_pp = True class-attribute ¶
A flag that indicates this model supports pipeline parallel.
Note
There is no need to redefine this flag if this class is in the MRO of your model class.
forward(input_ids, positions, *, intermediate_tensors) ¶
Accept IntermediateTensors when PP rank > 0.
Return IntermediateTensors only for the last PP rank.
Source code in vllm/model_executor/models/interfaces.py
SupportsTranscription ¶
Bases: Protocol
The interface required for all models that support transcription.
Methods:
-
get_generation_prompt–Get the prompt for the ASR model.
-
get_language_detection_prompt–Return a prompt that triggers language detection.
-
get_language_token_ids–Return token IDs that represent valid language tokens.
-
get_num_audio_tokens–Map from audio duration to number of audio tokens produced by the ASR
-
get_speech_to_text_config–Get the speech to text config for the ASR model.
-
get_streaming_post_processor_cls–Return a stateful post-processor class for streaming output deltas.
-
parse_diarized_transcript–Parse the model-specific diarized transcript format.
-
parse_language_detection_output–Parse the detected language from model output token IDs.
-
post_process_output–Post-process the raw model output text.
-
validate_language–Ensure the language specified in the transcription request
Attributes:
-
no_space_languages(set[str]) –Languages that don't need a space between words.
-
supports_diarized_transcription(bool) –Enables the
diarized_jsonresponse format for the model. -
supports_explicit_language_detection(bool) –Transcription models that require an explicit language detection step
-
supports_segment_timestamp(bool) –Enables the segment timestamp option for supported models by setting this to
True. -
supports_transcription_only(bool) –Transcription models can opt out of text generation by setting this to
Source code in vllm/model_executor/models/interfaces.py
1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 | |
no_space_languages = {'ja', 'zh'} class-attribute ¶
Languages that don't need a space between words. For example, Japanese (ja) and Chinese (zh) don't need a space between words.
supports_diarized_transcription = False class-attribute ¶
Enables the diarized_json response format for the model.
supports_explicit_language_detection = False class-attribute ¶
Transcription models that require an explicit language detection step (e.g. Whisper needs a separate forward pass to predict the language token) should set this to True and implement :meth:get_language_detection_prompt and :meth:parse_language_detection_output and :meth:get_language_token_ids.
supports_segment_timestamp = False class-attribute ¶
Enables the segment timestamp option for supported models by setting this to True.
supports_transcription_only = False class-attribute ¶
Transcription models can opt out of text generation by setting this to True.
get_generation_prompt(stt_params) classmethod ¶
Get the prompt for the ASR model. The model has control over the construction, as long as it returns a valid PromptType.
Source code in vllm/model_executor/models/interfaces.py
get_language_detection_prompt(audio, stt_config) classmethod ¶
Return a prompt that triggers language detection.
Only needs to be implemented when supports_explicit_language_detection is True.
Source code in vllm/model_executor/models/interfaces.py
get_language_token_ids(tokenizer) classmethod ¶
Return token IDs that represent valid language tokens.
Used to constrain language detection to only produce valid language tokens.
Only needs to be implemented when supports_explicit_language_detection is True.
Source code in vllm/model_executor/models/interfaces.py
get_num_audio_tokens(audio_duration_s, stt_config, model_config) classmethod ¶
Map from audio duration to number of audio tokens produced by the ASR model, without running a forward pass. This is used for estimating the amount of processing for this audio.
Source code in vllm/model_executor/models/interfaces.py
get_speech_to_text_config(model_config, task_type) classmethod ¶
Get the speech to text config for the ASR model.
get_streaming_post_processor_cls() classmethod ¶
Return a stateful post-processor class for streaming output deltas.
Each instance receives the next decoded text delta and whether the request output is final. It returns the cleaned delta that should be sent to the client.
Source code in vllm/model_executor/models/interfaces.py
parse_diarized_transcript(text) classmethod ¶
Parse the model-specific diarized transcript format.
Only models that set supports_diarized_transcription must override this method.
Source code in vllm/model_executor/models/interfaces.py
parse_language_detection_output(token_ids, tokenizer) classmethod ¶
Parse the detected language from model output token IDs.
Only needs to be implemented when supports_explicit_language_detection is True.
Source code in vllm/model_executor/models/interfaces.py
post_process_output(text) classmethod ¶
Post-process the raw model output text.
Some ASR models output structured formats (e.g., language tags, special tokens) that need to be stripped before returning to the user.
Parameters:
Returns:
-
str–Cleaned transcription text.
Source code in vllm/model_executor/models/interfaces.py
validate_language(language) classmethod ¶
Ensure the language specified in the transcription request is a valid ISO 639-1 language code. If the request language is valid, but not natively supported by the model, trigger a warning (but not an exception).
Source code in vllm/model_executor/models/interfaces.py
VllmModelForPooling ¶
Bases: VllmModel[T_co], Protocol[T_co]
The interface required for all pooling models in vLLM.
Attributes:
-
attn_type(AttnTypeStr) –Indicates the
-
default_seq_pooling_type(SequencePoolingType) –Indicates the vllm.config.pooler.PoolerConfig.seq_pooling_type
-
default_tok_pooling_type(TokenPoolingType) –Indicates the vllm.config.pooler.PoolerConfig.tok_pooling_type
-
is_pooling_model(Literal[True]) –A flag that indicates this model supports pooling.
-
pooler(Pooler) –The pooler is only called on TP rank 0.
-
score_type(ScoreType) –Indicates the
Source code in vllm/model_executor/models/interfaces_base.py
attn_type = 'decoder' class-attribute ¶
Indicates the vllm.config.model.ModelConfig.attn_type to use by default.
You can use the vllm.model_executor.models.interfaces_base.attn_type decorator to conveniently set this field.
default_seq_pooling_type = 'LAST' class-attribute ¶
Indicates the vllm.config.pooler.PoolerConfig.seq_pooling_type to use by default.
You can use the vllm.model_executor.models.interfaces_base.default_pooling_type decorator to conveniently set this field.
default_tok_pooling_type = 'ALL' class-attribute ¶
Indicates the vllm.config.pooler.PoolerConfig.tok_pooling_type to use by default.
You can use the vllm.model_executor.models.interfaces_base.default_pooling_type decorator to conveniently set this field.
is_pooling_model = True class-attribute ¶
A flag that indicates this model supports pooling.
Note
There is no need to redefine this flag if this class is in the MRO of your model class.
pooler instance-attribute ¶
The pooler is only called on TP rank 0.
score_type = 'bi-encoder' class-attribute ¶
Indicates the vllm.config.model.ModelConfig.score_type to use by default.
Scoring API handles score/rerank for:
-
"classify" task (score_type: cross-encoder models)
-
"embed" task (score_type: bi-encoder models)
-
"token_embed" task (score_type: late interaction models)
score_type defaults to bi-encoder, then the Score API uses the "embed" task.
If you set score_type to cross-encoder via vllm.model_executor.models.interfaces.SupportsCrossEncoding, then the Score API uses the "score" task.
If you set score_type to late-interaction via vllm.model_executor.models.interfaces.SupportsLateInteraction, then the Score API uses the "token_embed" task.
VllmModelForTextGeneration ¶
Bases: VllmModel[T], Protocol[T]
The interface required for all generative models in vLLM.
Methods:
-
compute_logits–Return
Noneif TP rank > 0.