[megatron] megatron remove num_samples by Jintao-Huang · Pull Request #9522 · modelscope/ms-swift

Jintao-Huang · 2026-06-09T08:00:12Z

No description provided.

gemini-code-assist

Code Review

This pull request simplifies batch preparation and data collation by removing the num_samples parameter and instead deriving it dynamically from the length of seq_lens when packed_seq_params is present. While this cleanup streamlines the code, the review highlights a critical regression in get_last_tokens where packed_seq_params.num_samples is still accessed. Additionally, the feedback suggests adding safety checks when popping seq_lens from the batch and when accessing it within packed_seq_params to prevent potential KeyError or AttributeError exceptions.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Jintao-Huang · 2026-06-15T08:14:51Z

/gemini review

gemini-code-assist

Code Review

This pull request refactors the tracking and calculation of num_samples across various Megatron trainers, primarily deriving it from packed_seq_params.seq_lens.shape[0] during padding-free training rather than passing it explicitly or storing it in batch dictionaries. It also introduces a new test file tests/megatron/test_opsd.py. A critical issue was identified in swift/megatron/trainers/rlhf_mixin.py where accessing packed_seq_params.seq_lens directly will cause an AttributeError if packed_seq_params is None (such as during standard training without packing). A code suggestion has been provided to safely fall back to labels.shape[0] in this scenario.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

gemini-code-assist · 2026-06-15T08:16:13Z

        per_token_logps = -output_tensor
        loss_mask = labels != -100
        per_token_logps = per_token_logps * loss_mask
+        num_samples = packed_seq_params.seq_lens.shape[0]


When padding_free is False, packed_seq_params is None. Accessing packed_seq_params.seq_lens will raise an AttributeError: 'NoneType' object has no attribute 'seq_lens', causing an unconditional crash during standard training (e.g., standard DPO/KTO without packing). We should safely fall back to labels.shape[0] when packed_seq_params is None.

Suggested change

num_samples = packed_seq_params.seq_lens.shape[0]

num_samples = packed_seq_params.seq_lens.shape[0] if packed_seq_params is not None else labels.shape[0]

Jintao-Huang · 2026-06-15T09:31:26Z

/gemini review

gemini-code-assist

Code Review

This pull request refactors the handling of num_samples across various Megatron trainers and utility functions. Instead of passing num_samples as an explicit parameter or retrieving it from batch dictionaries, it is now dynamically calculated using packed_seq_params.seq_lens.shape[0] or len(seq_lens). This simplifies function signatures and improves consistency. Additionally, test configurations are updated, and a new test file test_opsd.py is added. There are no review comments provided, so I have no feedback to address.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

megatron remove num_samples

de8b44f

gemini-code-assist Bot reviewed Jun 9, 2026

View reviewed changes

Comment thread swift/megatron/trainers/base.py

Comment thread swift/megatron/trainers/utils.py Outdated

Comment thread swift/megatron/trainers/utils.py Outdated

update

2cd141b

Jintao-Huang force-pushed the megatron_remove_num_samples branch from 823caf5 to 2cd141b Compare June 9, 2026 09:54

Jintao-Huang added 6 commits June 10, 2026 10:23

Merge branch 'main' into megatron_remove_num_samples

3d3ac36

Merge branch 'main' into megatron_remove_num_samples

fcd9248

Merge branch 'main' into megatron_remove_num_samples

2a3831d

update

3dc00cd

update

d27c40b

update

f6c74dd

gemini-code-assist Bot reviewed Jun 15, 2026

View reviewed changes

Jintao-Huang added 3 commits June 15, 2026 16:23

update

7e6570d

fix

e3f9b87

fix

4bc2474

hjh0119 approved these changes Jun 15, 2026

View reviewed changes

gemini-code-assist Bot reviewed Jun 15, 2026

View reviewed changes

Merge branch 'main' into megatron_remove_num_samples

91f6302

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[megatron] megatron remove num_samples#9522

[megatron] megatron remove num_samples#9522
Jintao-Huang wants to merge 12 commits into
modelscope:mainfrom
Jintao-Huang:megatron_remove_num_samples

Jintao-Huang commented Jun 9, 2026

Uh oh!

gemini-code-assist Bot left a comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Jintao-Huang commented Jun 15, 2026

Uh oh!

gemini-code-assist Bot left a comment

Uh oh!

gemini-code-assist Bot Jun 15, 2026

Uh oh!

Jintao-Huang commented Jun 15, 2026

Uh oh!

gemini-code-assist Bot left a comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

	num_samples = packed_seq_params.seq_lens.shape[0]
	num_samples = packed_seq_params.seq_lens.shape[0] if packed_seq_params is not None else labels.shape[0]

Conversation

Jintao-Huang commented Jun 9, 2026

Uh oh!

gemini-code-assist Bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Jintao-Huang commented Jun 15, 2026

Uh oh!

gemini-code-assist Bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

gemini-code-assist Bot Jun 15, 2026

Choose a reason for hiding this comment

Uh oh!

Jintao-Huang commented Jun 15, 2026

Uh oh!

gemini-code-assist Bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants