画像のアルファチャンネルをlossのマスクとして使用するオプションを追加 (#1223)

* Add alpha_mask parameter and apply masked loss * Fix type hint in trim_and_resize_if_required function * Refactor code to use keyword arguments in train_util.py * Fix alpha mask flipping logic * Fix alpha mask initialization * Fix alpha_mask transformation * Cache alpha_mask * Update alpha_masks to be on CPU * Set flipped_alpha_masks to Null if option disabled * Check if alpha_mask is None * Set alpha_mask to None if option disabled * Add description of alpha_mask option to docs
2026-04-06 13:47:06 +00:00 · 2024-05-19 19:07:25 +09:00
parent febc5c59fa
commit db6752901f
10 changed files with 105 additions and 129 deletions
--- a/train_textual_inversion_XTI.py
+++ b/train_textual_inversion_XTI.py
@@ -475,7 +475,9 @@ def train(args):

                loss = train_util.conditional_loss(noise_pred.float(), target.float(), reduction="none", loss_type=args.loss_type, huber_c=huber_c)
                if args.masked_loss:
-                    loss = apply_masked_loss(loss, batch)
+                    loss = apply_masked_loss(loss, batch["conditioning_images"][:, 0].unsqueeze(1))
+                if "alpha_mask" in batch and batch["alpha_mask"] is not None:
+                    loss = apply_masked_loss(loss, batch["alpha_mask"])
                loss = loss.mean([1, 2, 3])

                loss_weights = batch["loss_weights"]  # 各sampleごとのweight