feat: add pyramid noise and noise offset options to generation script (#2252)

* feat: add pyramid noise and noise offset options to generation script * fix: fix to work with SD1.5 models * doc: update to match with latest gen_img.py * doc: update README to clarify script capabilities and remove deprecated sections
2026-04-06 05:44:56 +00:00 · 2026-01-18 16:56:48 +09:00 · 2026-01-18 16:56:48 +09:00 · a9af52692a
commit a9af52692a
parent c6bc632ec6
3 changed files with 372 additions and 300 deletions
--- a/docs/gen_img_README-ja.md
+++ b/docs/gen_img_README-ja.md
@ -1,30 +1,24 @@
-SD 1.xおよび2.xのモデル、当リポジトリで学習したLoRA、ControlNet（v1.0のみ動作確認）などに対応した、Diffusersベースの推論（画像生成）スクリプトです。コマンドラインから用います。
+SD 1.x、2.x、およびSDXLのモデル、当リポジトリで学習したLoRA、ControlNet、ControlNet-LLLiteなどに対応した、独自の推論（画像生成）スクリプトです。コマンドラインから用います。

 # 概要

-* Diffusers (v0.10.2) ベースの推論（画像生成）スクリプト。
+* 独自の推論（画像生成）スクリプト。
 * SD 1.x、2.x (base/v-parameterization)、およびSDXLモデルに対応。
 * txt2img、img2img、inpaintingに対応。
 * 対話モード、およびファイルからのプロンプト読み込み、連続生成に対応。
 * プロンプト1行あたりの生成枚数を指定可能。
 * 全体の繰り返し回数を指定可能。
 * `fp16`だけでなく`bf16`にも対応。
-* xformersに対応し高速生成が可能。
-    * xformersにより省メモリ生成を行いますが、Automatic 1111氏のWeb UIほど最適化していないため、512*512の画像生成でおおむね6GB程度のVRAMを使用します。
+* xformers、SDPA（Scaled Dot-Product Attention）に対応。
 * プロンプトの225トークンへの拡張。ネガティブプロンプト、重みづけに対応。
-* Diffusersの各種samplerに対応（Web UIよりもsampler数は少ないです）。
+* Diffusersの各種samplerに対応。
 * Text Encoderのclip skip（最後からn番目の層の出力を用いる）に対応。
-* VAEの別途読み込み。
-* CLIP Guided Stable Diffusion、VGG16 Guided Stable Diffusion、Highres. fix、upscale対応。
-    * Highres. fixはWeb UIの実装を全く確認していない独自実装のため、出力結果は異なるかもしれません。
-* LoRA対応。適用率指定、複数LoRA同時利用、重みのマージに対応。
-    * Text EncoderとU-Netで別の適用率を指定することはできません。
-* Attention Coupleに対応。
-* ControlNet v1.0に対応。
+* VAEの別途読み込み、VAEのバッチ処理やスライスによる省メモリ化に対応。
+* Highres. fix（独自実装およびGradual Latent）、upscale対応。
+* LoRA、DyLoRA対応。適用率指定、複数LoRA同時利用、重みのマージに対応。
+* Attention Couple、Regional LoRAに対応。
+* ControlNet (v1.0/v1.1)、ControlNet-LLLiteに対応。
 * 途中でモデルを切り替えることはできませんが、バッチファイルを組むことで対応できます。
-* 個人的に欲しくなった機能をいろいろ追加。
-
-機能追加時にすべてのテストを行っているわけではないため、以前の機能に影響が出て一部機能が動かない可能性があります。何か問題があればお知らせください。

 # 基本的な使い方

@ -33,18 +27,20 @@ SD 1.xおよび2.xのモデル、当リポジトリで学習したLoRA、Control
 以下のように入力してください。

 ```batchfile
-python gen_img_diffusers.py --ckpt <モデル名> --outdir <画像出力先> --xformers --fp16 --interactive
+python gen_img.py --ckpt <モデル名> --outdir <画像出力先> --xformers --fp16 --interactive
 ```

 `--ckpt`オプションにモデル（Stable Diffusionのcheckpointファイル、またはDiffusersのモデルフォルダ）、`--outdir`オプションに画像の出力先フォルダを指定します。

-`--xformers`オプションでxformersの使用を指定します（xformersを使わない場合は外してください）。`--fp16`オプションでfp16（単精度）での推論を行います。RTX 30系のGPUでは `--bf16`オプションでbf16（bfloat16）での推論を行うこともできます。
+`--xformers`オプションでxformersの使用を指定します。`--fp16`オプションでfp16（半精度）での推論を行います。RTX 30系以降のGPUでは `--bf16`オプションでbf16（bfloat16）での推論を行うこともできます。

 `--interactive`オプションで対話モードを指定しています。

 Stable Diffusion 2.0（またはそこからの追加学習モデル）を使う場合は`--v2`オプションを追加してください。v-parameterizationを使うモデル（`768-v-ema.ckpt`およびそこからの追加学習モデル）を使う場合はさらに`--v_parameterization`を追加してください。

-`--v2`の指定有無が間違っているとモデル読み込み時にエラーになります。`--v_parameterization`の指定有無が間違っていると茶色い画像が表示されます。
+SDXLモデルを使う場合は`--sdxl`オプションを追加してください。
+
+`--v2`や`--sdxl`の指定有無が間違っているとモデル読み込み時にエラーになります。`--v_parameterization`の指定有無が間違っていると茶色い画像が表示されます。

 `Type prompt:`と表示されたらプロンプトを入力してください。

@ -59,7 +55,7 @@ Stable Diffusion 2.0（またはそこからの追加学習モデル）を使う
 以下のように入力します（実際には1行で入力します）。

 ```batchfile
-python gen_img_diffusers.py --ckpt <モデル名> --outdir <画像出力先> 
+python gen_img.py --ckpt <モデル名> --outdir <画像出力先> 
    --xformers --fp16 --images_per_prompt <生成枚数> --prompt "<プロンプト>"
 ```

@ -72,7 +68,7 @@ python gen_img_diffusers.py --ckpt <モデル名> --outdir <画像出力先>
 以下のように入力します。

 ```batchfile
-python gen_img_diffusers.py --ckpt <モデル名> --outdir <画像出力先> 
+python gen_img.py --ckpt <モデル名> --outdir <画像出力先> 
    --xformers --fp16 --from_file <プロンプトファイル名>
 ```

@ -106,7 +102,17 @@ python gen_img_diffusers.py --ckpt <モデル名> --outdir <画像出力先>
    
    `--v2`や`--sdxl`の指定有無が間違っているとモデル読み込み時にエラーになります。`--v_parameterization`の指定有無が間違っていると茶色い画像が表示されます。

- `--vae`：使用するVAEを指定します。未指定時はモデル内のVAEを使用します。
+- `--zero_terminal_snr`：noise schedulerのbetasを修正して、zero terminal SNRを強制します。
+
+- `--pyramid_noise_prob`：ピラミッドノイズを適用する確率を指定します。
+
+- `--pyramid_noise_discount_range`：ピラミッドノイズの割引率の範囲を指定します。
+
+- `--noise_offset_prob`：ノイズオフセットを適用する確率を指定します。
+
+- `--noise_offset_range`：ノイズオフセットの範囲を指定します。
+
+- `--vae`：使用する VAE を指定します。未指定時はモデル内の VAE を使用します。

 - `--tokenizer_cache_dir`：トークナイザーのキャッシュディレクトリを指定します（オフライン利用のため）。

@ -130,13 +136,14 @@ python gen_img_diffusers.py --ckpt <モデル名> --outdir <画像出力先>

 - `--scale <ガイダンススケール>`：unconditionalガイダンススケールを指定します。デフォルトは`7.5`です。

- `--sampler <サンプラー名>`：サンプラーを指定します。デフォルトは`ddim`です。Diffusersで提供されているddim、pndm、dpmsolver、dpmsolver+++、lms、euler、euler_a、が指定可能です（後ろの三つはk_lms、k_euler、k_euler_aでも指定できます）。
+- `--sampler <サンプラー名>`：サンプラーを指定します。デフォルトは`ddim`です。
+    `ddim`, `pndm`, `lms`, `euler`, `euler_a`, `heun`, `dpm_2`, `dpm_2_a`, `dpmsolver`, `dpmsolver++`, `dpmsingle`, `k_lms`, `k_euler`, `k_euler_a`, `k_dpm_2`, `k_dpm_2_a` が指定可能です。

 - `--outdir <画像出力先フォルダ>`：画像の出力先を指定します。

 - `--images_per_prompt <生成枚数>`：プロンプト1件当たりの生成枚数を指定します。デフォルトは`1`です。

- `--clip_skip <スキップ数>`：CLIPの後ろから何番目の層を使うかを指定します。省略時は最後の層を使います。
+- `--clip_skip <スキップ数>`：CLIPの後ろから何番目の層を使うかを指定します。デフォルトはSD1/2の場合1、SDXLの場合2です。

 - `--max_embeddings_multiples <倍数>`：CLIPの入出力長をデフォルト（75）の何倍にするかを指定します。未指定時は75のままです。たとえば3を指定すると入出力長が225になります。

@ -144,6 +151,8 @@ python gen_img_diffusers.py --ckpt <モデル名> --outdir <画像出力先>

 - `--emb_normalize_mode`：embedding正規化モードを指定します。"original"（デフォルト）、"abs"、"none"から選択できます。プロンプトの重みの正規化方法に影響します。

+- `--force_scheduler_zero_steps_offset`：スケジューラのステップオフセットを、スケジューラ設定の `steps_offset` の値に関わらず強制的にゼロにします。
+
 ## SDXL固有のオプション

 SDXL モデル（`--sdxl`フラグ付き）を使用する場合、追加のコンディショニングオプションが利用できます：
@ -164,7 +173,7 @@ SDXL モデル（`--sdxl`フラグ付き）を使用する場合、追加のコ

 - `--batch_size <バッチサイズ>`：バッチサイズを指定します。デフォルトは`1`です。バッチサイズが大きいとメモリを多く消費しますが、生成速度が速くなります。

- `--vae_batch_size <VAEのバッチサイズ>`：VAEのバッチサイズを指定します。デフォルトはバッチサイズと同じです。
+- `--vae_batch_size <VAEのバッチサイズ>`：VAEのバッチサイズを指定します。デフォルトはバッチサイズと同じです。1未満の値を指定すると、バッチサイズに対する比率として扱われます。
    VAEのほうがメモリを多く消費するため、デノイジング後（stepが100%になった後）でメモリ不足になる場合があります。このような場合にはVAEのバッチサイズを小さくしてください。

 - `--vae_slices <スライス数>`：VAE処理時に画像をスライスに分割してVRAM使用量を削減します。None（デフォルト）で分割なし。16や32のような値が推奨されます。有効にすると処理が遅くなりますが、VRAM使用量が少なくなります。
@ -177,9 +186,9 @@ SDXL モデル（`--sdxl`フラグ付き）を使用する場合、追加のコ

 - `--diffusers_xformers`：Diffusers経由でxformersを使用します（注：Hypernetworksと互換性がありません）。

- `--fp16`：fp16（単精度）での推論を行います。`fp16`と`bf16`をどちらも指定しない場合はfp32（単精度）での推論を行います。
+- `--fp16`：fp16（半精度）での推論を行います。`fp16`と`bf16`をどちらも指定しない場合はfp32（単精度）での推論を行います。

- `--bf16`：bf16（bfloat16）での推論を行います。RTX 30系のGPUでのみ指定可能です。`--bf16`オプションはRTX 30系以外のGPUではエラーになります。`fp16`よりも`bf16`のほうが推論結果がNaNになる（真っ黒の画像になる）可能性が低いようです。
+- `--bf16`：bf16（bfloat16）での推論を行います。RTX 30系以降のGPUでのみ指定可能です。`--bf16`オプションはRTX 30系以外のGPUではエラーになります。SDXLでは`fp16`よりも`bf16`のほうが推論結果がNaNになる（真っ黒の画像になる）可能性が低いようです。

 ## 追加ネットワーク（LoRA等）の使用

@ -204,7 +213,7 @@ SDXL モデル（`--sdxl`フラグ付き）を使用する場合、追加のコ
 次は同一プロンプトで64枚をバッチサイズ4で一括生成する例です。

 ```batchfile
-python gen_img_diffusers.py --ckpt model.ckpt --outdir outputs 
+python gen_img.py --ckpt model.ckpt --outdir outputs 
    --xformers --fp16 --W 512 --H 704 --scale 12.5 --sampler k_euler_a 
    --steps 32 --batch_size 4 --images_per_prompt 64 
    --prompt "beautiful flowers --n monochrome"
@ -213,7 +222,7 @@ python gen_img_diffusers.py --ckpt model.ckpt --outdir outputs
 次はファイルに書かれたプロンプトを、それぞれ10枚ずつ、バッチサイズ4で一括生成する例です。

 ```batchfile
-python gen_img_diffusers.py --ckpt model.ckpt --outdir outputs 
+python gen_img.py --ckpt model.ckpt --outdir outputs 
    --xformers --fp16 --W 512 --H 704 --scale 12.5 --sampler k_euler_a 
    --steps 32 --batch_size 4 --images_per_prompt 10 
    --from_file prompts.txt
@ -222,7 +231,7 @@ python gen_img_diffusers.py --ckpt model.ckpt --outdir outputs
 Textual Inversion（後述）およびLoRAの使用例です。

 ```batchfile
-python gen_img_diffusers.py --ckpt model.safetensors 
+python gen_img.py --ckpt model.safetensors 
    --scale 8 --steps 48 --outdir txt2img --xformers 
    --W 512 --H 768 --fp16 --sampler k_euler_a 
    --textual_inversion_embeddings goodembed.safetensors negprompt.pt 
@ -258,6 +267,22 @@ python gen_img_diffusers.py --ckpt model.safetensors

 - `--am`：追加ネットワークの重みを指定します。コマンドラインからの指定を上書きします。複数の追加ネットワークを使用する場合は`--am 0.8,0.5,0.3`のように __カンマ区切りで__ 指定します。

+- `--ow`：SDXLのoriginal_widthを指定します。
+
+- `--oh`：SDXLのoriginal_heightを指定します。
+
+- `--nw`：SDXLのoriginal_width_negativeを指定します。
+
+- `--nh`：SDXLのoriginal_height_negativeを指定します。
+
+- `--ct`：SDXLのcrop_topを指定します。
+
+- `--cl`：SDXLのcrop_leftを指定します。
+
+- `--c`：CLIPプロンプトを指定します。
+
+- `--f`：生成ファイル名を指定します。
+
 ※これらのオプションを指定すると、バッチサイズよりも小さいサイズでバッチが実行される場合があります（これらの値が異なると一括生成できないため）。（あまり気にしなくて大丈夫ですが、ファイルからプロンプトを読み込み生成する場合は、これらの値が同一のプロンプトを並べておくと効率が良くなります。）

 例：
@ -267,6 +292,21 @@ python gen_img_diffusers.py --ckpt model.safetensors

 ![image](https://user-images.githubusercontent.com/52813779/235343446-25654172-fff4-4aaf-977a-20d262b51676.png)

+# プロンプトのワイルドカード (Dynamic Prompts)
+
+Dynamic Prompts (Wildcard) 記法に対応しています。Web UIの拡張機能等と完全に同じではありませんが、以下の機能が利用可能です。
+
+- `{A|B|C}` : A, B, C の中からランダムに1つを選択します。
+- `{e$$A|B|C}` : A, B, C のすべてを順に利用します（全列挙）。プロンプト内に複数の `{e$$...}` がある場合、すべての組み合わせが生成されます。
+  - 例：`{e$$red|blue} flower, {e$$1girl|2girls}` → `red flower, 1girl`, `red flower, 2girls`, `blue flower, 1girl`, `blue flower, 2girls` の4枚が生成されます。
+- `{n$$A|B|C}` : A, B, C の中から n 個をランダムに選択して結合します。
+  - 例：`{2$$A|B|C}` → `A, B` や `B, C` など。
+- `{n-m$$A|B|C}` : A, B, C の中から n 個から m 個をランダムに選択して結合します。
+- `{$$sep$$A|B|C}` : 選択された項目を sep で結合します（デフォルトは `, `）。
+  - 例：`{2$$ and $$A|B|C}` → `A and B` など。
+
+これらは組み合わせて利用可能です。
+
 # img2img

 ## オプション
@ -284,7 +324,7 @@ python gen_img_diffusers.py --ckpt model.safetensors
 ## コマンドラインからの実行例

 ```batchfile
-python gen_img_diffusers.py --ckpt trinart_characters_it4_v1_vae_merged.ckpt 
+python gen_img.py --ckpt trinart_characters_it4_v1_vae_merged.ckpt 
    --outdir outputs --xformers --fp16 --scale 12.5 --sampler k_euler --steps 32 
    --image_path template.png --strength 0.8 
    --prompt "1girl, cowboy shot, brown hair, pony tail, brown eyes, 
@ -325,10 +365,6 @@ img2img時にコマンドラインオプションの`--W`と`--H`で生成画像

 モデルとして、当リポジトリで学習したTextual Inversionモデル、およびWeb UIで学習したTextual Inversionモデル（画像埋め込みは非対応）を利用できます

-## Extended Textual Inversion
-
-`--textual_inversion_embeddings`の代わりに`--XTI_embeddings`オプションを指定してください。使用法は`--textual_inversion_embeddings`と同じです。
-
 ## Highres. fix

 AUTOMATIC1111氏のWeb UIにある機能の類似機能です（独自実装のためもしかしたらいろいろ異なるかもしれません）。最初に小さめの画像を生成し、その画像を元にimg2imgすることで、画像全体の破綻を防ぎつつ大きな解像度の画像を生成します。
@ -343,6 +379,8 @@ img2imgと併用できません。

 - `--highres_fix_steps`：1st stageの画像のステップ数を指定します。デフォルトは`28`です。

+- `--highres_fix_strength`：1st stageのimg2img時のstrengthを指定します。省略時は`--strength`と同じ値になります。
+
 - `--highres_fix_save_1st`：1st stageの画像を保存するかどうかを指定します。

 - `--highres_fix_latents_upscaling`：指定すると2nd stageの画像生成時に1st stageの画像をlatentベースでupscalingします（bilinearのみ対応）。未指定時は画像をLANCZOS4でupscalingします。
@ -357,7 +395,7 @@ img2imgと併用できません。
 コマンドラインの例です。

 ```batchfile
-python gen_img_diffusers.py  --ckpt trinart_characters_it4_v1_vae_merged.ckpt
+python gen_img.py  --ckpt trinart_characters_it4_v1_vae_merged.ckpt
    --n_iter 1 --scale 7.5 --W 1024 --H 1024 --batch_size 1 --outdir ../txt2img 
    --steps 48 --sampler ddim --fp16 
    --xformers 
@ -407,16 +445,16 @@ Deep Shrinkは、異なるタイムステップで異なる深度のUNetを使
 - `--control_net_preps`：ControlNetのプリプロセスを指定します。`--control_net_models`と同様に複数指定可能です。現在はcannyのみ対応しています。対象モデルでプリプロセスを使用しない場合は `none` を指定します。
   cannyの場合 `--control_net_preps canny_63_191`のように、閾値1と2を'_'で区切って指定できます。

- `--control_net_weights`：ControlNetの適用時の重みを指定します（`1.0`で通常、`0.5`なら半分の影響力で適用）。`--control_net_models`と同様に複数指定可能です。
+- `--control_net_multipliers`：ControlNetの適用時の重みを指定します（`1.0`で通常、`0.5`なら半分の影響力で適用）。`--control_net_models`と同様に複数指定可能です。

 - `--control_net_ratios`：ControlNetを適用するstepの範囲を指定します。`0.5`の場合は、step数の半分までControlNetを適用します。`--control_net_models`と同様に複数指定可能です。

 コマンドラインの例です。

 ```batchfile
-python gen_img_diffusers.py --ckpt model_ckpt --scale 8 --steps 48 --outdir txt2img --xformers 
+python gen_img.py --ckpt model_ckpt --scale 8 --steps 48 --outdir txt2img --xformers 
    --W 512 --H 768 --bf16 --sampler k_euler_a 
-    --control_net_models diff_control_sd15_canny.safetensors --control_net_weights 1.0 
+    --control_net_models diff_control_sd15_canny.safetensors --control_net_multipliers 1.0 
    --guide_image_path guide.png --control_net_ratios 1.0 --interactive
 ```

@ -458,70 +496,6 @@ ControlNetと組み合わせることも可能です（細かい位置指定に

 LoRAを指定すると、`--network_weights`で指定した複数のLoRAがそれぞれANDの各部分に対応します。現在の制約として、LoRAの数はANDの部分の数と同じである必要があります。

-## CLIP Guided Stable Diffusion
-
-DiffusersのCommunity Examplesの[こちらのcustom pipeline](https://github.com/huggingface/diffusers/blob/main/examples/community/README.md#clip-guided-stable-diffusion)からソースをコピー、変更したものです。
-
-通常のプロンプトによる生成指定に加えて、追加でより大規模のCLIPでプロンプトのテキストの特徴量を取得し、生成中の画像の特徴量がそのテキストの特徴量に近づくよう、生成される画像をコントロールします（私のざっくりとした理解です）。大きめのCLIPを使いますのでVRAM使用量はかなり増加し（VRAM 8GBでは512*512でも厳しいかもしれません）、生成時間も掛かります。
-
-なお選択できるサンプラーはDDIM、PNDM、LMSのみとなります。
-
-`--clip_guidance_scale`オプションにどの程度、CLIPの特徴量を反映するかを数値で指定します。先のサンプルでは100になっていますので、そのあたりから始めて増減すると良いようです。
-
-デフォルトではプロンプトの先頭75トークン（重みづけの特殊文字を除く）がCLIPに渡されます。プロンプトの`--c`オプションで、通常のプロンプトではなく、CLIPに渡すテキストを別に指定できます（たとえばCLIPはDreamBoothのidentifier（識別子）や「1girl」などのモデル特有の単語は認識できないと思われますので、それらを省いたテキストが良いと思われます）。
-
-コマンドラインの例です。
-
-```batchfile
-python gen_img_diffusers.py  --ckpt v1-5-pruned-emaonly.ckpt --n_iter 1 
-    --scale 2.5 --W 512 --H 512 --batch_size 1 --outdir ../txt2img --steps 36  
-    --sampler ddim --fp16 --opt_channels_last --xformers --images_per_prompt 1  
-    --interactive --clip_guidance_scale 100
-```
-
-## CLIP Image Guided Stable Diffusion
-
-テキストではなくCLIPに別の画像を渡し、その特徴量に近づくよう生成をコントロールする機能です。`--clip_image_guidance_scale`オプションで適用量の数値を、`--guide_image_path`オプションでguideに使用する画像（ファイルまたはフォルダ）を指定してください。
-
-コマンドラインの例です。
-
-```batchfile
-python gen_img_diffusers.py  --ckpt trinart_characters_it4_v1_vae_merged.ckpt
-    --n_iter 1 --scale 7.5 --W 512 --H 512 --batch_size 1 --outdir ../txt2img 
-    --steps 80 --sampler ddim --fp16 --opt_channels_last --xformers 
-    --images_per_prompt 1  --interactive  --clip_image_guidance_scale 100 
-    --guide_image_path YUKA160113420I9A4104_TP_V.jpg
-```
-
-### VGG16 Guided Stable Diffusion
-
-指定した画像に近づくように画像生成する機能です。通常のプロンプトによる生成指定に加えて、追加でVGG16の特徴量を取得し、生成中の画像が指定したガイド画像に近づくよう、生成される画像をコントロールします。img2imgでの使用をお勧めします（通常の生成では画像がぼやけた感じになります）。CLIP Guided Stable Diffusionの仕組みを流用した独自の機能です。またアイデアはVGGを利用したスタイル変換から拝借しています。
-
-なお選択できるサンプラーはDDIM、PNDM、LMSのみとなります。
-
-`--vgg16_guidance_scale`オプションにどの程度、VGG16特徴量を反映するかを数値で指定します。試した感じでは100くらいから始めて増減すると良いようです。`--guide_image_path`オプションでguideに使用する画像（ファイルまたはフォルダ）を指定してください。
-
-複数枚の画像を一括でimg2img変換し、元画像をガイド画像とする場合、`--guide_image_path`と`--image_path`に同じ値を指定すればOKです。
-
-コマンドラインの例です。
-
-```batchfile
-python gen_img_diffusers.py --ckpt wd-v1-3-full-pruned-half.ckpt 
-    --n_iter 1 --scale 5.5 --steps 60 --outdir ../txt2img 
-    --xformers --sampler ddim --fp16 --W 512 --H 704 
-    --batch_size 1 --images_per_prompt 1 
-    --prompt "picturesque, 1girl, solo, anime face, skirt, beautiful face 
-        --n lowres, bad anatomy, bad hands, error, missing fingers, 
-        cropped, worst quality, low quality, normal quality, 
-        jpeg artifacts, blurry, 3d, bad face, monochrome --d 1" 
-    --strength 0.8 --image_path ..\src_image
-    --vgg16_guidance_scale 100 --guide_image_path ..\src_image 
-```
-
-`--vgg16_guidance_layerPで特徴量取得に使用するVGG16のレイヤー番号を指定できます（デフォルトは20でconv4-2のReLUです）。上の層ほど画風を表現し、下の層ほどコンテンツを表現するといわれています。
-
-![image](https://user-images.githubusercontent.com/52813779/235343813-3c1f0d7a-4fb3-4274-98e4-b92d76b551df.png)
-
 # その他のオプション

 - `--no_preview` : 対話モードでプレビュー画像を表示しません。OpenCVがインストールされていない場合や、出力されたファイルを直接確認する場合に指定してください。
@ -542,27 +516,11 @@ python gen_img_diffusers.py --ckpt wd-v1-3-full-pruned-half.ckpt

 - `--network_show_meta`：追加ネットワークのメタデータを表示します。

-
 --- 

-# About Gradual Latent
-
-Gradual Latent is a Hires fix that gradually increases the size of the latent.  `gen_img.py`, `sdxl_gen_img.py`, and `gen_img_diffusers.py` have the following options.
-
- `--gradual_latent_timesteps`: Specifies the timestep to start increasing the size of the latent. The default is None, which means Gradual Latent is not used. Please try around 750 at first.
- `--gradual_latent_ratio`: Specifies the initial size of the latent. The default is 0.5, which means it starts with half the default latent size.
- `--gradual_latent_ratio_step`: Specifies the ratio to increase the size of the latent. The default is 0.125, which means the latent size is gradually increased to 0.625, 0.75, 0.875, 1.0.
- `--gradual_latent_ratio_every_n_steps`: Specifies the interval to increase the size of the latent. The default is 3, which means the latent size is increased every 3 steps.
-
-Each option can also be specified with prompt options, `--glt`, `--glr`, `--gls`, `--gle`.
-
-__Please specify `euler_a` for the sampler.__ Because the source code of the sampler is modified. It will not work with other samplers.
-
-It is more effective with SD 1.5. It is quite subtle with SDXL.
-
 # Gradual Latent について

-latentのサイズを徐々に大きくしていくHires fixです。`gen_img.py` 、``sdxl_gen_img.py` 、`gen_img_diffusers.py` に以下のオプションが追加されています。
+latentのサイズを徐々に大きくしていくHires fixです。

 - `--gradual_latent_timesteps` : latentのサイズを大きくし始めるタイムステップを指定します。デフォルトは None で、Gradual Latentを使用しません。750 くらいから始めてみてください。
 - `--gradual_latent_ratio` : latentの初期サイズを指定します。デフォルトは 0.5 で、デフォルトの latent サイズの半分のサイズから始めます。
--- a/docs/gen_img_README.md
+++ b/docs/gen_img_README.md
@ -10,25 +10,16 @@ This is an inference (image generation) script that supports SD 1.x and 2.x mode
 * The number of images generated per prompt line can be specified.
 * The total number of repetitions can be specified.
 * Supports not only `fp16` but also `bf16`.
-* Supports xformers for high-speed generation.
-    * Although xformers are used for memory-saving generation, it is not as optimized as Automatic 1111's Web UI, so it uses about 6GB of VRAM for 512*512 image generation.
+* Supports xformers and SDPA (Scaled Dot-Product Attention).
 * Extension of prompts to 225 tokens. Supports negative prompts and weighting.
-* Supports various samplers from Diffusers including ddim, pndm, lms, euler, euler_a, heun, dpm_2, dpm_2_a, dpmsolver, dpmsolver++, dpmsingle.
+* Supports various samplers from Diffusers.
 * Supports clip skip (uses the output of the nth layer from the end) of Text Encoder.
-* Separate loading of VAE.
-* Supports CLIP Guided Stable Diffusion, VGG16 Guided Stable Diffusion, Highres. fix, and upscale.
-    * Highres. fix is an original implementation that has not confirmed the Web UI implementation at all, so the output results may differ.
-* LoRA support. Supports application rate specification, simultaneous use of multiple LoRAs, and weight merging.
-    * It is not possible to specify different application rates for Text Encoder and U-Net.
-* Supports Attention Couple.
-* Supports ControlNet v1.0.
-* Supports Deep Shrink for optimizing generation at different depths.
-* Supports Gradual Latent for progressive upscaling during generation.
-* Supports CLIP Vision Conditioning for img2img.
+* Separate loading of VAE, supports VAE batch processing and slicing for memory saving.
+* Highres. fix (original implementation and Gradual Latent), upscale support.
+* LoRA, DyLoRA support. Supports application rate specification, simultaneous use of multiple LoRAs, and weight merging.
+* Supports Attention Couple, Regional LoRA.
+* Supports ControlNet (v1.0/v1.1), ControlNet-LLLite.
 * It is not possible to switch models midway, but it can be handled by creating a batch file.
-* Various personally desired features have been added.
-
-Since not all tests are performed when adding features, it is possible that previous features may be affected and some features may not work. Please let us know if you have any problems.

 # Basic Usage

@ -110,6 +101,16 @@ Specify from the command line.

    If the `--v2` or `--sdxl` specification is incorrect, an error will occur when loading the model. If the `--v_parameterization` specification is incorrect, a brown image will be displayed.

+- `--zero_terminal_snr`: Modifies the noise scheduler betas to enforce zero terminal SNR.
+
+- `--pyramid_noise_prob`: Specifies the probability of applying pyramid noise.
+
+- `--pyramid_noise_discount_range`: Specifies the discount range for pyramid noise.
+
+- `--noise_offset_prob`: Specifies the probability of applying noise offset.
+
+- `--noise_offset_range`: Specifies the range of noise offset.
+
 - `--vae`: Specifies the VAE to use. If not specified, the VAE in the model will be used.

 - `--tokenizer_cache_dir`: Specifies the cache directory for the tokenizer (for offline usage).
@ -134,13 +135,14 @@ Specify from the command line.

 - `--scale <guidance_scale>`: Specifies the unconditional guidance scale. The default is `7.5`.

- `--sampler <sampler_name>`: Specifies the sampler. The default is `ddim`. The following samplers are supported: ddim, pndm, lms, euler, euler_a, heun, dpm_2, dpm_2_a, dpmsolver, dpmsolver++, dpmsingle. Some can also be specified with k_ prefix (k_lms, k_euler, k_euler_a, k_dpm_2, k_dpm_2_a).
+- `--sampler <sampler_name>`: Specifies the sampler. The default is `ddim`.
+    `ddim`, `pndm`, `lms`, `euler`, `euler_a`, `heun`, `dpm_2`, `dpm_2_a`, `dpmsolver`, `dpmsolver++`, `dpmsingle`, `k_lms`, `k_euler`, `k_euler_a`, `k_dpm_2`, `k_dpm_2_a` can be specified.

 - `--outdir <image_output_destination_folder>`: Specifies the output destination for images.

 - `--images_per_prompt <number_of_images_to_generate>`: Specifies the number of images to generate per prompt. The default is `1`.

- `--clip_skip <number_of_skips>`: Specifies which layer from the end of CLIP to use. If omitted, the last layer is used.
+- `--clip_skip <number_of_skips>`: Specifies which layer from the end of CLIP to use. Default is 1 for SD1/2, 2 for SDXL.

 - `--max_embeddings_multiples <multiplier>`: Specifies how many times the CLIP input/output length should be multiplied by the default (75). If not specified, it remains 75. For example, specifying 3 makes the input/output length 225.

@ -148,6 +150,8 @@ Specify from the command line.

 - `--emb_normalize_mode`: Specifies the embedding normalization mode. Options are "original" (default), "abs", and "none". This affects how prompt weights are normalized.

+- `--force_scheduler_zero_steps_offset`: Forces the scheduler step offset to zero regardless of the `steps_offset` value in the scheduler configuration.
+
 ## SDXL-Specific Options

 When using SDXL models (with `--sdxl` flag), additional conditioning options are available:
@ -262,6 +266,22 @@ Please put spaces before and after the prompt option specification `--n`.

 - `--am`: Specifies the weight of the additional network. Overrides the command line specification. If using multiple additional networks, specify them separated by __commas__, like `--am 0.8,0.5,0.3`.

+- `--ow`: Specifies original_width for SDXL.
+
+- `--oh`: Specifies original_height for SDXL.
+
+- `--nw`: Specifies original_width_negative for SDXL.
+
+- `--nh`: Specifies original_height_negative for SDXL.
+
+- `--ct`: Specifies crop_top for SDXL.
+
+- `--cl`: Specifies crop_left for SDXL.
+
+- `--c`: Specifies the CLIP prompt.
+
+- `--f`: Specifies the generated file name.
+
 - `--glt`: Specifies the timestep to start increasing the size of the latent for Gradual Latent. Overrides the command line specification.

 - `--glr`: Specifies the initial size of the latent for Gradual Latent as a ratio. Overrides the command line specification.
@ -279,6 +299,21 @@ Example:

 ![image](https://user-images.githubusercontent.com/52813779/235343446-25654172-fff4-4aaf-977a-20d262b51676.png)

+# Wildcards in Prompts (Dynamic Prompts)
+
+Dynamic Prompts (Wildcard) notation is supported. While not exactly the same as the Web UI extension, the following features are available.
+
+- `{A|B|C}` : Randomly selects one from A, B, or C.
+- `{e$$A|B|C}` : Uses all of A, B, and C in order (enumeration). If there are multiple `{e$$...}` in the prompt, all combinations will be generated.
+  - Example: `{e$$red|blue} flower, {e$$1girl|2girls}` -> Generates 4 images: `red flower, 1girl`, `red flower, 2girls`, `blue flower, 1girl`, `blue flower, 2girls`.
+- `{n$$A|B|C}` : Randomly selects n items from A, B, C and combines them.
+  - Example: `{2$$A|B|C}` -> `A, B` or `B, C`, etc.
+- `{n-m$$A|B|C}` : Randomly selects between n and m items from A, B, C and combines them.
+- `{$$sep$$A|B|C}` : Combines selected items with `sep` (default is `, `).
+  - Example: `{2$$ and $$A|B|C}` -> `A and B`, etc.
+
+These can be used in combination.
+
 # img2img

 ## Options
@ -337,10 +372,6 @@ Specify the embeddings to use with the `--textual_inversion_embeddings` option (

 As models, you can use Textual Inversion models trained with this repository and Textual Inversion models trained with Web UI (image embedding is not supported).

-## Extended Textual Inversion
-
-Specify the `--XTI_embeddings` option instead of `--textual_inversion_embeddings`. Usage is the same as `--textual_inversion_embeddings`.
-
 ## Highres. fix

 This is a similar feature to the one in AUTOMATIC1111's Web UI (it may differ in various ways as it is an original implementation). It first generates a smaller image and then uses that image as a base for img2img to generate a large resolution image while preventing the entire image from collapsing.
@ -480,70 +511,6 @@ It can also be combined with ControlNet (combination with ControlNet is recommen

 If LoRA is specified, multiple LoRAs specified with `--network_weights` will correspond to each part of AND. As a current constraint, the number of LoRAs must be the same as the number of AND parts.

-## CLIP Guided Stable Diffusion
-
-The source code is copied and modified from [this custom pipeline](https://github.com/huggingface/diffusers/blob/main/examples/community/README.md#clip-guided-stable-diffusion) in Diffusers' Community Examples.
-
-In addition to the normal prompt-based generation specification, it additionally acquires the text features of the prompt with a larger CLIP and controls the generated image so that the features of the image being generated approach those text features (this is my rough understanding). Since a larger CLIP is used, VRAM usage increases considerably (it may be difficult even for 512*512 with 8GB of VRAM), and generation time also increases.
-
-Note that the selectable samplers are DDIM, PNDM, and LMS only.
-
-Specify how much to reflect the CLIP features numerically with the `--clip_guidance_scale` option. In the previous sample, it is 100, so it seems good to start around there and increase or decrease it.
-
-By default, the first 75 tokens of the prompt (excluding special weighting characters) are passed to CLIP. With the `--c` option in the prompt, you can specify the text to be passed to CLIP separately from the normal prompt (for example, it is thought that CLIP cannot recognize DreamBooth identifiers or model-specific words like "1girl", so text excluding them is considered good).
-
-Command line example:
-
-```batchfile
-python gen_img.py  --ckpt v1-5-pruned-emaonly.ckpt --n_iter 1 \
-    --scale 2.5 --W 512 --H 512 --batch_size 1 --outdir ../txt2img --steps 36  \
-    --sampler ddim --fp16 --opt_channels_last --xformers --images_per_prompt 1  \
-    --interactive --clip_guidance_scale 100
-```
-
-## CLIP Image Guided Stable Diffusion
-
-This is a feature that passes another image to CLIP instead of text and controls generation to approach its features. Specify the numerical value of the application amount with the `--clip_image_guidance_scale` option and the image (file or folder) to use for guidance with the `--guide_image_path` option.
-
-Command line example:
-
-```batchfile
-python gen_img.py  --ckpt trinart_characters_it4_v1_vae_merged.ckpt\
-    --n_iter 1 --scale 7.5 --W 512 --H 512 --batch_size 1 --outdir ../txt2img \
-    --steps 80 --sampler ddim --fp16 --opt_channels_last --xformers \
-    --images_per_prompt 1  --interactive  --clip_image_guidance_scale 100 \
-    --guide_image_path YUKA160113420I9A4104_TP_V.jpg
-```
-
-### VGG16 Guided Stable Diffusion
-
-This is a feature that generates images to approach a specified image. In addition to the normal prompt-based generation specification, it additionally acquires the features of VGG16 and controls the generated image so that the image being generated approaches the specified guide image. It is recommended to use it with img2img (images tend to be blurred in normal generation). This is an original feature that reuses the mechanism of CLIP Guided Stable Diffusion. The idea is also borrowed from style transfer using VGG.
-
-Note that the selectable samplers are DDIM, PNDM, and LMS only.
-
-Specify how much to reflect the VGG16 features numerically with the `--vgg16_guidance_scale` option. From what I've tried, it seems good to start around 100 and increase or decrease it. Specify the image (file or folder) to use for guidance with the `--guide_image_path` option.
-
-When batch converting multiple images with img2img and using the original images as guide images, it is OK to specify the same value for `--guide_image_path` and `--image_path`.
-
-Command line example:
-
-```batchfile
-python gen_img.py --ckpt wd-v1-3-full-pruned-half.ckpt \
-    --n_iter 1 --scale 5.5 --steps 60 --outdir ../txt2img \
-    --xformers --sampler ddim --fp16 --W 512 --H 704 \
-    --batch_size 1 --images_per_prompt 1 \
-    --prompt "picturesque, 1girl, solo, anime face, skirt, beautiful face \
-        --n lowres, bad anatomy, bad hands, error, missing fingers, \
-        cropped, worst quality, low quality, normal quality, \
-        jpeg artifacts, blurry, 3d, bad face, monochrome --d 1" \
-    --strength 0.8 --image_path ..\\src_image\
-    --vgg16_guidance_scale 100 --guide_image_path ..\\src_image \
-```
-
-You can specify the VGG16 layer number used for feature acquisition with `--vgg16_guidance_layerP` (default is 20, which is ReLU of conv4-2). It is said that upper layers express style and lower layers express content.
-
-![image](https://user-images.githubusercontent.com/52813779/235343813-3c1f0d7a-4fb3-4274-98e4-b92d76b551df.png)
-
 # Other Options

 - `--no_preview`: Does not display preview images in interactive mode. Specify this if OpenCV is not installed or if you want to check the output files directly.
@ -576,7 +543,7 @@ Gradual Latent is a Hires fix that gradually increases the size of the latent.
 - `--gradual_latent_ratio_step`: Specifies the ratio to increase the size of the latent. The default is 0.125, which means the latent size is gradually increased to 0.625, 0.75, 0.875, 1.0.
 - `--gradual_latent_ratio_every_n_steps`: Specifies the interval to increase the size of the latent. The default is 3, which means the latent size is increased every 3 steps.
 - `--gradual_latent_s_noise`: Specifies the s_noise parameter for Gradual Latent. Default is 1.0.
- `--gradual_latent_unsharp_params`: Specifies unsharp mask parameters for Gradual Latent in the format: ksize,sigma,strength,target-x (where target-x: 1=True, 0=False). Recommended values: `3,0.5,0.5,1` or `3,1.0,1.0,0`.
+- `--gradual_latent_unsharp_params`: Specifies unsharp mask parameters for Gradual Latent in the format: ksize,sigma,strength,target-x (target-x: 1=True, 0=False). Recommended values: `3,0.5,0.5,1` or `3,1.0,1.0,0`.

 Each option can also be specified with prompt options, `--glt`, `--glr`, `--gls`, `--gle`.

--- a/gen_img.py
+++ b/gen_img.py
@ -1,5 +1,6 @@
 import itertools
 import json
+from types import SimpleNamespace
 from typing import Any, List, NamedTuple, Optional, Tuple, Union, Callable
 import glob
 import importlib
@ -20,7 +21,8 @@ import diffusers
 import numpy as np
 import torch

-from library.device_utils import init_ipex, clean_memory, get_preferred_device
+from library.device_utils import init_ipex
+from library.strategy_sd import SdTokenizeStrategy

 init_ipex()

@ -60,6 +62,7 @@ from library.original_unet import UNet2DConditionModel, InferUNet2DConditionMode
 from library.sdxl_original_unet import InferSdxlUNet2DConditionModel
 from library.sdxl_original_control_net import SdxlControlNet
 from library.original_unet import FlashAttentionFunction
+from library.custom_train_functions import pyramid_noise_like
 from networks.control_net_lllite import ControlNetLLLite
 from library.utils import GradualLatent, EulerAncestralDiscreteSchedulerGL
 from library.utils import setup_logging, add_logging_arguments
@ -434,6 +437,7 @@ class PipelineLike:
        img2img_noise=None,
        clip_guide_images=None,
        emb_normalize_mode: str = "original",
+        force_scheduler_zero_steps_offset: bool = False,
        **kwargs,
    ):
        # TODO support secondary prompt
@ -707,7 +711,10 @@ class PipelineLike:
                    raise ValueError("The mask and init_image should be the same size!")

            # get the original timestep using init_timestep
-            offset = self.scheduler.config.get("steps_offset", 0)
+            if force_scheduler_zero_steps_offset:
+                offset = 0
+            else:
+                offset = self.scheduler.config.get("steps_offset", 0)
            init_timestep = int(num_inference_steps * strength) + offset
            init_timestep = min(init_timestep, num_inference_steps)

@ -859,7 +866,7 @@ class PipelineLike:
                            )
                        input_resi_add = input_resi_add_mean
                        mid_add = torch.mean(torch.stack(mid_add_list), dim=0)
-                        
+
                    noise_pred = self.unet(latent_model_input, t, text_embeddings, vector_embeddings, input_resi_add, mid_add)
            elif self.is_sdxl:
                noise_pred = self.unet(latent_model_input, t, text_embeddings, vector_embeddings)
@ -1362,97 +1369,177 @@ def preprocess_mask(mask):
 RE_DYNAMIC_PROMPT = re.compile(r"\{((e|E)\$\$)?(([\d\-]+)\$\$)?(([^\|\}]+?)\$\$)?(.+?((\|).+?)*?)\}")


-def handle_dynamic_prompt_variants(prompt, repeat_count):
+def handle_dynamic_prompt_variants(prompt, repeat_count, seed_random, seeds=None):
    founds = list(RE_DYNAMIC_PROMPT.finditer(prompt))
    if not founds:
-        return [prompt]
+        return [prompt], seeds

-    # make each replacement for each variant
-    enumerating = False
-    replacers = []
-    for found in founds:
-        # if "e$$" is found, enumerate all variants
-        found_enumerating = found.group(2) is not None
-        enumerating = enumerating or found_enumerating
+    # Prepare seeds list
+    if seeds is None:
+        seeds = []
+    while len(seeds) < repeat_count:
+        seeds.append(seed_random.randint(0, 2**32 - 1))

-        separator = ", " if found.group(6) is None else found.group(6)
-        variants = found.group(7).split("|")
+    # Escape braces
+    prompt = prompt.replace(r"\{", "｛").replace(r"\}", "｝")

-        # parse count range
-        count_range = found.group(4)
-        if count_range is None:
-            count_range = [1, 1]
-        else:
-            count_range = count_range.split("-")
-            if len(count_range) == 1:
-                count_range = [int(count_range[0]), int(count_range[0])]
-            elif len(count_range) == 2:
-                count_range = [int(count_range[0]), int(count_range[1])]
+    # Process nested dynamic prompts recursively
+    prompts = [prompt] * repeat_count
+    has_dynamic = True
+    while has_dynamic:
+        has_dynamic = False
+        new_prompts = []
+        for i, prompt in enumerate(prompts):
+            seed = seeds[i] if i < len(seeds) else seeds[0]  # if enumerating, use the first seed
+
+            # find innermost dynamic prompts
+
+            # find outer dynamic prompt and temporarily replace them with placeholders
+            deepest_nest_level = 0
+            nest_level = 0
+            for c in prompt:
+                if c == "{":
+                    nest_level += 1
+                    deepest_nest_level = max(deepest_nest_level, nest_level)
+                elif c == "}":
+                    nest_level -= 1
+            if deepest_nest_level == 0:
+                new_prompts.append(prompt)
+                continue  # no more dynamic prompts
+
+            # find positions of innermost dynamic prompts
+            positions = []
+            nest_level = 0
+            start_pos = -1
+            for i, c in enumerate(prompt):
+                if c == "{":
+                    nest_level += 1
+                    if nest_level == deepest_nest_level:
+                        start_pos = i
+                elif c == "}":
+                    if nest_level == deepest_nest_level:
+                        end_pos = i + 1
+                        positions.append((start_pos, end_pos))
+                    nest_level -= 1
+
+            # extract innermost dynamic prompts
+            innermost_founds = []
+            for start, end in positions:
+                segment = prompt[start:end]
+                m = RE_DYNAMIC_PROMPT.match(segment)
+                if m:
+                    innermost_founds.append((m, start, end))
+
+            if not innermost_founds:
+                new_prompts.append(prompt)
+                continue
+            has_dynamic = True
+
+            # make each replacement for each variant
+            enumerating = False
+            replacers = []
+            for found, start, end in innermost_founds:
+                # if "e$$" is found, enumerate all variants
+                found_enumerating = found.group(2) is not None
+                enumerating = enumerating or found_enumerating
+
+                separator = ", " if found.group(6) is None else found.group(6)
+                variants = found.group(7).split("|")
+
+                # parse count range
+                count_range = found.group(4)
+                if count_range is None:
+                    count_range = [1, 1]
+                else:
+                    count_range = count_range.split("-")
+                    if len(count_range) == 1:
+                        count_range = [int(count_range[0]), int(count_range[0])]
+                    elif len(count_range) == 2:
+                        count_range = [int(count_range[0]), int(count_range[1])]
+                    else:
+                        logger.warning(f"invalid count range: {count_range}")
+                        count_range = [1, 1]
+                    if count_range[0] > count_range[1]:
+                        count_range = [count_range[1], count_range[0]]
+                    if count_range[0] < 0:
+                        count_range[0] = 0
+                    if count_range[1] > len(variants):
+                        count_range[1] = len(variants)
+
+                if found_enumerating:
+                    # make function to enumerate all combinations
+                    def make_replacer_enum(vari, cr, sep):
+                        def replacer(rnd=random):
+                            values = []
+                            for count in range(cr[0], cr[1] + 1):
+                                for comb in itertools.combinations(vari, count):
+                                    values.append(sep.join(comb))
+                            return values
+
+                        return replacer
+
+                    replacers.append(make_replacer_enum(variants, count_range, separator))
+                else:
+                    # make function to choose random combinations
+                    def make_replacer_single(vari, cr, sep):
+                        def replacer(rnd=random):
+                            count = rnd.randint(cr[0], cr[1])
+                            comb = rnd.sample(vari, count)
+                            return [sep.join(comb)]
+
+                        return replacer
+
+                    replacers.append(make_replacer_single(variants, count_range, separator))
+
+            # make each prompt
+            rnd = random.Random(seed)
+            if not enumerating:
+                # if not enumerating, repeat the prompt, replace each variant randomly
+
+                # reverse the lists to replace from end to start, keep positions correct
+                innermost_founds.reverse()
+                replacers.reverse()
+
+                current = prompt
+                for (found, start, end), replacer in zip(innermost_founds, replacers):
+                    current = current[:start] + replacer(rnd)[0] + current[end:]
+                new_prompts.append(current)
            else:
-                logger.warning(f"invalid count range: {count_range}")
-                count_range = [1, 1]
-            if count_range[0] > count_range[1]:
-                count_range = [count_range[1], count_range[0]]
-            if count_range[0] < 0:
-                count_range[0] = 0
-            if count_range[1] > len(variants):
-                count_range[1] = len(variants)
+                # if enumerating, iterate all combinations for previous prompts, all seeds are same
+                processing_prompts = [prompt]
+                for found, replacer in zip(founds, replacers):
+                    if found.group(2) is not None:
+                        # make all combinations for existing prompts
+                        repleced_prompts = []
+                        for current in processing_prompts:
+                            replacements = replacer(rnd)
+                            for replacement in replacements:
+                                repleced_prompts.append(
+                                    current.replace(found.group(0), replacement, 1)
+                                )  # This does not work if found is duplicated
+                        processing_prompts = repleced_prompts

-        if found_enumerating:
-            # make function to enumerate all combinations
-            def make_replacer_enum(vari, cr, sep):
-                def replacer():
-                    values = []
-                    for count in range(cr[0], cr[1] + 1):
-                        for comb in itertools.combinations(vari, count):
-                            values.append(sep.join(comb))
-                    return values
+                for found, replacer in zip(founds, replacers):
+                    # make random selection for existing prompts
+                    if found.group(2) is None:
+                        for i in range(len(processing_prompts)):
+                            processing_prompts[i] = processing_prompts[i].replace(found.group(0), replacer(rnd)[0], 1)

-                return replacer
+                new_prompts.extend(processing_prompts)

-            replacers.append(make_replacer_enum(variants, count_range, separator))
-        else:
-            # make function to choose random combinations
-            def make_replacer_single(vari, cr, sep):
-                def replacer():
-                    count = random.randint(cr[0], cr[1])
-                    comb = random.sample(vari, count)
-                    return [sep.join(comb)]
+        prompts = new_prompts

-                return replacer
+    # Restore escaped braces
+    for i in range(len(prompts)):
+        prompts[i] = prompts[i].replace("｛", "{").replace("｝", "}")
+    if enumerating:
+        # adjust seeds list
+        new_seeds = []
+        for _ in range(len(prompts)):
+            new_seeds.append(seeds[0])  # use the first seed for all
+        seeds = new_seeds

-            replacers.append(make_replacer_single(variants, count_range, separator))
-
-    # make each prompt
-    if not enumerating:
-        # if not enumerating, repeat the prompt, replace each variant randomly
-        prompts = []
-        for _ in range(repeat_count):
-            current = prompt
-            for found, replacer in zip(founds, replacers):
-                current = current.replace(found.group(0), replacer()[0], 1)
-            prompts.append(current)
-    else:
-        # if enumerating, iterate all combinations for previous prompts
-        prompts = [prompt]
-
-        for found, replacer in zip(founds, replacers):
-            if found.group(2) is not None:
-                # make all combinations for existing prompts
-                new_prompts = []
-                for current in prompts:
-                    replecements = replacer()
-                    for replecement in replecements:
-                        new_prompts.append(current.replace(found.group(0), replecement, 1))
-                prompts = new_prompts
-
-        for found, replacer in zip(founds, replacers):
-            # make random selection for existing prompts
-            if found.group(2) is None:
-                for i in range(len(prompts)):
-                    prompts[i] = prompts[i].replace(found.group(0), replacer()[0], 1)
-
-    return prompts
+    return prompts, seeds


 # endregion
@ -1612,7 +1699,8 @@ def main(args):
        tokenizers = [tokenizer1, tokenizer2]
    else:
        if use_stable_diffusion_format:
-            tokenizer = train_util.load_tokenizer(args)
+            tokenize_strategy = SdTokenizeStrategy(args.v2, max_length=None, tokenizer_cache_dir=args.tokenizer_cache_dir)
+            tokenizer = tokenize_strategy.tokenizer
        tokenizers = [tokenizer]

    # schedulerを用意する
@ -1719,6 +1807,9 @@ def main(args):
    if scheduler_module is not None:
        scheduler_module.torch = TorchRandReplacer(noise_manager)

+    if args.zero_terminal_snr:
+        sched_init_args["rescale_betas_zero_snr"] = True
+
    scheduler = scheduler_cls(
        num_train_timesteps=SCHEDULER_TIMESTEPS,
        beta_start=SCHEDULER_LINEAR_START,
@ -1727,6 +1818,9 @@ def main(args):
        **sched_init_args,
    )

+    # if args.zero_terminal_snr:
+    #     custom_train_functions.fix_noise_scheduler_betas_for_zero_terminal_snr(scheduler)
+
    # ↓以下は結局PipeでFalseに設定されるので意味がなかった
    # # clip_sample=Trueにする
    # if hasattr(scheduler.config, "clip_sample") and scheduler.config.clip_sample is False:
@ -1868,7 +1962,7 @@ def main(args):
        if not is_sdxl:
            for i, model in enumerate(args.control_net_models):
                prep_type = None if not args.control_net_preps or len(args.control_net_preps) <= i else args.control_net_preps[i]
-                weight = 1.0 if not args.control_net_weights or len(args.control_net_weights) <= i else args.control_net_weights[i]
+                weight = 1.0 if not args.control_net_multipliers or len(args.control_net_multipliers) <= i else args.control_net_multipliers[i]
                ratio = 1.0 if not args.control_net_ratios or len(args.control_net_ratios) <= i else args.control_net_ratios[i]

                ctrl_unet, ctrl_net = original_control_net.load_control_net(args.v2, unet, model)
@ -2355,7 +2449,9 @@ def main(args):
                    if images_1st.dtype == torch.bfloat16:
                        images_1st = images_1st.to(torch.float)  # interpolateがbf16をサポートしていない
                    images_1st = torch.nn.functional.interpolate(
-                        images_1st, (batch[0].ext.height // 8, batch[0].ext.width // 8), mode="bilinear"
+                        images_1st,
+                        (batch[0].ext.height // 8, batch[0].ext.width // 8),
+                        mode="bicubic",
                    )  # , antialias=True)
                    images_1st = images_1st.to(org_dtype)

@ -2464,6 +2560,20 @@ def main(args):
                torch.manual_seed(seed)
                start_code[i] = torch.randn(noise_shape, device=device, dtype=dtype)

+                # pyramid noise
+                if args.pyramid_noise_prob is not None and random.random() < args.pyramid_noise_prob:
+                    min_discount, max_discount = args.pyramid_noise_discount_range
+                    discount = torch.rand(1, device=device, dtype=dtype) * (max_discount - min_discount) + min_discount
+                    logger.info(f"apply pyramid noise to start code: {start_code[i].shape}, discount: {discount.item()}")
+                    start_code[i] = pyramid_noise_like(start_code[i].unsqueeze(0), device=device, discount=discount).squeeze(0)
+
+                # noise offset
+                if args.noise_offset_prob is not None and random.random() < args.noise_offset_prob:
+                    min_offset, max_offset = args.noise_offset_range
+                    noise_offset = torch.randn(1, device=device, dtype=dtype) * (max_offset - min_offset) + min_offset
+                    logger.info(f"apply noise offset to start code: {start_code[i].shape}, offset: {noise_offset.item()}")
+                    start_code[i] += noise_offset
+
                # make each noises
                for j in range(steps * scheduler_num_noises_per_step):
                    noises[j][i] = torch.randn(noise_shape, device=device, dtype=dtype)
@ -2532,6 +2642,7 @@ def main(args):
                clip_prompts=clip_prompts,
                clip_guide_images=guide_images,
                emb_normalize_mode=args.emb_normalize_mode,
+                force_scheduler_zero_steps_offset=args.force_scheduler_zero_steps_offset,
            )
            if highres_1st and not args.highres_fix_save_1st:  # return images or latents
                return images
@ -2624,7 +2735,16 @@ def main(args):

            # sd-dynamic-prompts like variants:
            # count is 1 (not dynamic) or images_per_prompt (no enumeration) or arbitrary (enumeration)
-            raw_prompts = handle_dynamic_prompt_variants(raw_prompt, args.images_per_prompt)
+            seeds = None
+            m = re.search(r" --d ([\d+,]+)", raw_prompt, re.IGNORECASE)
+            if m:
+                seeds = [int(d) for d in m[0][5:].split(",")]
+                logger.info(f"seeds: {seeds}")
+                raw_prompt = raw_prompt[: m.start()] + raw_prompt[m.end() :]
+
+            raw_prompts, prompt_seeds = handle_dynamic_prompt_variants(raw_prompt, args.images_per_prompt, seed_random, seeds)
+            if prompt_seeds is not None:
+                seeds = prompt_seeds

            # repeat prompt
            for pi in range(args.images_per_prompt if len(raw_prompts) == 1 else len(raw_prompts)):
@ -2644,8 +2764,8 @@ def main(args):
                    scale = args.scale
                    negative_scale = args.negative_scale
                    steps = args.steps
-                    seed = None
-                    seeds = None
+                    # seed = None
+                    # seeds = None
                    strength = 0.8 if args.strength is None else args.strength
                    negative_prompt = ""
                    clip_prompt = None
@ -2727,11 +2847,11 @@ def main(args):
                                logger.info(f"steps: {steps}")
                                continue

-                            m = re.match(r"d ([\d,]+)", parg, re.IGNORECASE)
-                            if m:  # seed
-                                seeds = [int(d) for d in m.group(1).split(",")]
-                                logger.info(f"seeds: {seeds}")
-                                continue
+                            # m = re.match(r"d ([\d,]+)", parg, re.IGNORECASE)
+                            # if m:  # seed
+                            #     seeds = [int(d) for d in m.group(1).split(",")]
+                            #     logger.info(f"seeds: {seeds}")
+                            #     continue

                            m = re.match(r"l ([\d\.]+)", parg, re.IGNORECASE)
                            if m:  # scale
@ -3012,6 +3132,27 @@ def setup_parser() -> argparse.ArgumentParser:
    parser.add_argument(
        "--v_parameterization", action="store_true", help="enable v-parameterization training / v-parameterization学習を有効にする"
    )
+    parser.add_argument(
+        "--zero_terminal_snr",
+        action="store_true",
+        help="fix noise scheduler betas to enforce zero terminal SNR / noise schedulerのbetasを修正して、zero terminal SNRを強制する",
+    )
+    parser.add_argument(
+        "--pyramid_noise_prob", type=float, default=None, help="probability for pyramid noise / ピラミッドノイズの確率"
+    )
+    parser.add_argument(
+        "--pyramid_noise_discount_range",
+        type=float,
+        nargs=2,
+        default=None,
+        help="discount range for pyramid noise / ピラミッドノイズの割引範囲",
+    )
+    parser.add_argument(
+        "--noise_offset_prob", type=float, default=None, help="probability for noise offset / ノイズオフセットの確率"
+    )
+    parser.add_argument(
+        "--noise_offset_range", type=float, nargs=2, default=None, help="range for noise offset / ノイズオフセットの範囲"
+    )

    parser.add_argument("--prompt", type=str, default=None, help="prompt / プロンプト")
    parser.add_argument(
@ -3250,6 +3391,12 @@ def setup_parser() -> argparse.ArgumentParser:
        choices=["original", "none", "abs"],
        help="embedding normalization mode / embeddingの正規化モード",
    )
+    parser.add_argument(
+        "--force_scheduler_zero_steps_offset",
+        action="store_true",
+        help="force scheduler steps offset to zero"
+        + " / スケジューラのステップオフセットをスケジューラ設定の `steps_offset` の値に関わらず強制的にゼロにする",
+    )
    parser.add_argument(
        "--guide_image_path", type=str, default=None, nargs="*", help="image to ControlNet / ControlNetでガイドに使う画像"
    )