> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://fr.nvidia-localization.ferndocs.com/fr/nemo/rl/latest/nemo-rl-models-generation-vllm-vllm-worker-async/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://fr.nvidia-localization.ferndocs.com/_mcp/server. # nemo\_rl.models.generation.vllm.vllm\_worker\_async ## Contenu du module ### Classes | Nom | Description | | -------------------------------------------------------------------------------------------------- | ----------- | | [`VllmAsyncGenerationWorker`](#nemorlmodelsgenerationvllmvllmworkerasyncvllmasyncgenerationworker) | Aucun | ### API ```python class nemo_rl.models.generation.vllm.vllm_worker_async.VllmAsyncGenerationWorker ``` **Bases** : `nemo_rl.models.generation.vllm.vllm_worker.BaseVllmGenerationWorker` ```python _create_engine(llm_kwargs: dict[str, typing.Any]) -> None ``` ```python post_init_async() ``` ```python init_collective_async( rank_prefix: int, ip: str, port: int, world_size: int ) -> None ``` ```python generate_async( data: nemo_rl.distributed.batched_data_dict.BatchedDataDict[nemo_rl.models.generation.interfaces.GenerationDatumSpec], greedy: bool = False ) -> typing.AsyncGenerator[tuple[int, nemo_rl.distributed.batched_data_dict.BatchedDataDict[nemo_rl.models.generation.interfaces.GenerationOutputSpec]], None] ``` Générer un lot de données à l'aide du moteur AsyncLLM de vLLM, en renvoyant les résultats dès qu'ils sont prêts. Args : data : BatchedDataDict avec input\_ids et input\_lengths greedy : Indique s'il faut utiliser le décodage glouton au lieu de l'échantillonnage Renvoie : Tuple de (indice\_original, BatchedDataDict conforme à GenerationOutputSpec pour la séquence unique) ```python generate_text_async( data: nemo_rl.distributed.batched_data_dict.BatchedDataDict[nemo_rl.models.generation.interfaces.GenerationDatumSpec], greedy: bool = False ) -> typing.AsyncGenerator[tuple[int, nemo_rl.distributed.batched_data_dict.BatchedDataDict[nemo_rl.models.generation.interfaces.GenerationOutputSpec]], None] ``` Générer de façon asynchrone des réponses textuelles, en renvoyant les résultats dès qu'ils sont prêts. Args : data : BatchedDataDict contenant des invites avec des chaînes de texte greedy : Indique s'il faut utiliser le décodage glouton au lieu de l'échantillonnage Renvoie : Tuple de (indice\_original, BatchedDataDict contenant une seule réponse textuelle) ```python report_device_id_async() -> list[str] ``` Version asynchrone de report\_device\_id. ```python prepare_refit_info_async(state_dict_info: dict[str, typing.Any]) -> None ``` Version asynchrone de prepare\_refit\_info. ```python update_weights_from_ipc_handles_async(ipc_handles: dict[str, typing.Any]) -> bool ``` Version asynchrone de update\_weights\_from\_ipc\_handles. Args : ipc\_handles (dict) : Dictionnaire mappant les UUID de périphériques (str) aux handles IPC de paramètres. Renvoie : bool : True si les poids ont été mis à jour avec succès, False sinon. ```python update_weights_from_collective_async() -> bool ``` Version asynchrone de update\_weights\_from\_collective. ```python reset_prefix_cache_async() ``` Version asynchrone de reset\_prefix\_cache. ```python sleep_async() ``` Version asynchrone de sleep. ```python wake_up_async(**kwargs) ``` Version asynchrone de wake\_up. ```python shutdown() -> bool ``` Nettoyer les ressources vLLM.