Hermes Agent Wiki 非公式・日本語wiki
この skill をそのまま使う: GitHub で原文を見る

英語原文・frontmatter 込みで、Hermes が読み込む実体そのままです(このページの本文は日本語版)。

Tensorrt Llm

目次

NVIDIA の GPU で、LLM の推論を高い処理量で動かします。

skill の情報

提供元 追加インストール — hermes skills install official/mlops/tensorrt-llm で導入します
パス optional-skills/mlops/tensorrt-llm
バージョン 1.0.1
作者 Orchestra Research
ライセンス MIT
依存関係 tensorrt-llm, torch
対応プラットフォーム linux, macos
タグ Inference Serving, TensorRT-LLM, NVIDIA, Inference Optimization, High Throughput, Low Latency, Production, FP8, INT4, In-Flight Batching, Multi-GPU

参考: SKILL.md 全文

TensorRT-LLM

NVIDIA の GPU 上で LLM の推論を高い性能で動かすための、NVIDIA 製のオープンソースライブラリです。

TensorRT-LLM が向いているとき

次のようなときに使います:

  • NVIDIA の GPU(A100、H100、GB200)で動かす
  • 処理量を最大にしたい(Llama 3 で毎秒 24,000 トークン以上)
  • 応答が返るまでの遅れを小さくしたい
  • 量子化したモデル(FP8、INT4、FP4)を扱う
  • 複数の GPU やノードにまたがって動かす

代わりに vLLM を使うとき:

  • 準備を簡単に済ませ、Python 中心の書き方をしたい
  • TensorRT のコンパイルなしで PagedAttention を使いたい
  • AMD の GPU など、NVIDIA 以外のハードウェアで動かす

代わりに llama.cpp を使うとき:

  • CPU や Apple Silicon で動かす
  • NVIDIA の GPU がない環境で、端末側で動かしたい
  • GGUF という手軽な量子化形式を使いたい

すぐ試す

導入

# Docker (recommended) — images are on NGC (nvcr.io), not Docker Hub.
# Replace x.y.z with the desired version (e.g. 1.2.1). Browse tags on NGC:
# https://catalog.ngc.nvidia.com/orgs/nvidia/teams/tensorrt-llm/containers/release/tags
docker pull nvcr.io/nvidia/tensorrt-llm/release:x.y.z

# pip install (current stable GA)
pip install tensorrt_llm

# Requires CUDA 13.2.1, TensorRT 10.x, Python 3.10-3.12

まずは推論を動かす

from tensorrt_llm import LLM, SamplingParams

# Initialize model
llm = LLM(model="meta-llama/Meta-Llama-3-8B")

# Configure sampling
sampling_params = SamplingParams(
    max_tokens=100,
    temperature=0.7,
    top_p=0.9
)

# Generate
prompts = ["Explain quantum computing"]
outputs = llm.generate(prompts, sampling_params)

for output in outputs:
    print(output.text)

trtllm-serve で公開する

# Start server (automatic model download and compilation)
trtllm-serve meta-llama/Meta-Llama-3-8B \
    --tp_size 4 \              # Tensor parallelism (4 GPUs)
    --max_batch_size 256 \
    --max_num_tokens 4096

# Client request
curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B",
    "messages": [{"role": "user", "content": "Hello!"}],
    "temperature": 0.7,
    "max_tokens": 100
  }'

主な機能

性能まわりの工夫

  • In-flight batching: 生成の途中でも動的にまとめて処理します
  • Paged KV cache: メモリを無駄なく使います
  • Flash Attention: アテンションの計算を最適化します
  • 量子化: FP8、INT4、FP4 で推論が 2〜4 倍速くなります
  • CUDA graphs: カーネル起動の手間を減らします

並列化

  • テンソル並列(TP): モデルを GPU 間で分けます
  • パイプライン並列(PP): 層ごとに分けて配置します
  • エキスパート並列: Mixture-of-Experts のモデル向けです
  • 複数ノード: 1 台の枠を超えて広げられます

進んだ機能

  • 投機的デコード: 下書き用モデルを併用して生成を速めます
  • LoRA の提供: 複数のアダプタを効率よく配信します
  • 役割を分けた提供: 入力処理と生成を別々に動かします

よく使う書き方

量子化したモデル(FP8)

from tensorrt_llm import LLM

# Load FP8 quantized model (2× faster, 50% memory)
llm = LLM(
    model="meta-llama/Meta-Llama-3-70B",
    dtype="fp8",
    max_num_tokens=8192
)

# Inference same as before
outputs = llm.generate(["Summarize this article..."])

複数 GPU での運用

# Tensor parallelism across 8 GPUs
llm = LLM(
    model="meta-llama/Meta-Llama-3-405B",
    tensor_parallel_size=8,
    dtype="fp8"
)

まとめて推論する

# Process 100 prompts efficiently
prompts = [f"Question {i}: ..." for i in range(100)]

outputs = llm.generate(
    prompts,
    sampling_params=SamplingParams(max_tokens=200)
)

# Automatic in-flight batching for maximum throughput

性能の実測値

Meta Llama 3-8B(H100 GPU):

  • 処理量: 毎秒 24,000 トークン
  • 遅延: 1 トークンあたり約 10ms
  • PyTorch との比較: 100 倍の速さ

Llama 3-70B(A100 80GB × 8):

  • FP8 の量子化: FP16 の 2 倍の速さ
  • メモリ: FP8 で 50% 減ります

対応しているモデル

  • LLaMA 系: Llama 2、Llama 3、CodeLlama
  • GPT 系: GPT-2、GPT-J、GPT-NeoX
  • Qwen: Qwen、Qwen2、QwQ
  • DeepSeek: DeepSeek-V2、DeepSeek-V3
  • Mixtral: Mixtral-8x7B、Mixtral-8x22B
  • 画像対応: LLaVA、Phi-3-vision
  • HuggingFace 上に 100 以上のモデル

参考ドキュメント

参考情報

  • ドキュメント: https://nvidia.github.io/TensorRT-LLM/
  • GitHub: https://github.com/NVIDIA/TensorRT-LLM
  • モデル: https://huggingface.co/models?library=tensorrt_llm