> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://fr.nvidia-localization.ferndocs.com/fr/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://fr.nvidia-localization.ferndocs.com/fr/_mcp/server.

# nemo\_rl.models.generation.vllm.vllm\_worker\_async

## Contenu du module

### Classes

| Nom                                                                                                | Description |
| -------------------------------------------------------------------------------------------------- | ----------- |
| [`VllmAsyncGenerationWorker`](#nemorlmodelsgenerationvllmvllmworkerasyncvllmasyncgenerationworker) | Aucun       |

### API

```python
class nemo_rl.models.generation.vllm.vllm_worker_async.VllmAsyncGenerationWorker
```

**Bases** : `nemo_rl.models.generation.vllm.vllm_worker.BaseVllmGenerationWorker`

```python
_create_engine(llm_kwargs: dict[str, typing.Any]) -> None
```

```python
post_init_async()
```

```python
init_collective_async(
    rank_prefix: int,
    ip: str,
    port: int,
    world_size: int
) -> None
```

```python
generate_async(
    data: nemo_rl.distributed.batched_data_dict.BatchedDataDict[nemo_rl.models.generation.interfaces.GenerationDatumSpec],
    greedy: bool = False
) -> typing.AsyncGenerator[tuple[int, nemo_rl.distributed.batched_data_dict.BatchedDataDict[nemo_rl.models.generation.interfaces.GenerationOutputSpec]], None]
```

Générer un lot de données à l'aide du moteur AsyncLLM de vLLM, en renvoyant les résultats dès qu'ils sont prêts.

Args :
data : BatchedDataDict avec input\_ids et input\_lengths
greedy : Indique s'il faut utiliser le décodage glouton au lieu de l'échantillonnage

Renvoie :
Tuple de (indice\_original, BatchedDataDict conforme à GenerationOutputSpec pour la séquence unique)

```python
generate_text_async(
    data: nemo_rl.distributed.batched_data_dict.BatchedDataDict[nemo_rl.models.generation.interfaces.GenerationDatumSpec],
    greedy: bool = False
) -> typing.AsyncGenerator[tuple[int, nemo_rl.distributed.batched_data_dict.BatchedDataDict[nemo_rl.models.generation.interfaces.GenerationOutputSpec]], None]
```

Générer de façon asynchrone des réponses textuelles, en renvoyant les résultats dès qu'ils sont prêts.

Args :
data : BatchedDataDict contenant des invites avec des chaînes de texte
greedy : Indique s'il faut utiliser le décodage glouton au lieu de l'échantillonnage

Renvoie :
Tuple de (indice\_original, BatchedDataDict contenant une seule réponse textuelle)

```python
report_device_id_async() -> list[str]
```

Version asynchrone de report\_device\_id.

```python
prepare_refit_info_async(state_dict_info: dict[str, typing.Any]) -> None
```

Version asynchrone de prepare\_refit\_info.

```python
update_weights_from_ipc_handles_async(ipc_handles: dict[str, typing.Any]) -> bool
```

Version asynchrone de update\_weights\_from\_ipc\_handles.

Args :
ipc\_handles (dict) : Dictionnaire mappant les UUID de périphériques (str) aux handles IPC de paramètres.

Renvoie :
bool : True si les poids ont été mis à jour avec succès, False sinon.

```python
update_weights_from_collective_async() -> bool
```

Version asynchrone de update\_weights\_from\_collective.

```python
reset_prefix_cache_async()
```

Version asynchrone de reset\_prefix\_cache.

```python
sleep_async()
```

Version asynchrone de sleep.

```python
wake_up_async(**kwargs)
```

Version asynchrone de wake\_up.

```python
shutdown() -> bool
```

Nettoyer les ressources vLLM.