Add Autoconfig, Coordinated_Optimizer and Sharding keras implementations for Tensor Parallel Autosharding #21707

buildwithsuhana · 2025-10-02T10:19:24Z

This PR introduces support for tensor parallelism autosharding in Keras, enabling users to shard large model layers across multiple devices. This is a crucial feature for training models that are too large to fit into the memory of a single accelerator.

The implementation is centered around two new components:

autoconfig.py: This module contains the logic to analyze a Keras model, identify sharding candidates (e.g., Dense, EinsumDense layers), and generate a sharding plan.

coordinated_optimizer.py: This is an optimizer wrapper that consumes the sharding plan. During training, it intercepts gradients for sharded variables and performs a collective AllReduce to ensure weight updates are correctly synchronized across all devices.

Example usage: https://colab.research.google.com/drive/1UAINIcstDuO0aeA9lxCF5LaIj5ne5X5z?resourcekey=0-pPF4COO19KRoqS5cpWNILA&usp=sharing

This is the 2nd (out of 4) PR for AutoSharding Keras.

…hsuhana/keras into Tensor_parallel_keras_2

gemini-code-assist · 2025-10-02T10:19:41Z

Summary of Changes

Hello @buildwithsuhana, I'm Gemini Code Assist¹! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances Keras's capabilities for large-scale model training by introducing foundational support for tensor parallelism autosharding. It provides mechanisms to automatically determine how model layers should be split across multiple devices and a specialized optimizer to manage the distributed training process, including sharding optimizer states and synchronizing gradients. This enables users to train models that exceed the memory capacity of a single accelerator, making distributed training more accessible and efficient within the Keras ecosystem.

Highlights

Automatic Sharding Configuration: Introduced "autoconfig.py" to intelligently analyze Keras models (e.g., Dense, EinsumDense, Embedding layers) and generate a tensor parallelism sharding plan, classifying Dense layers as "up-projection" or "down-projection" for optimal splitting.
Coordinated Optimizer for Distributed Training: Added "coordinated_optimizer.py" which provides "CoordinatedOptimizer" for managing sharded optimizer states and synchronizing gradients across devices, and "TensorParallelOptimizer" as a Keras-compatible wrapper.
Gradient Synchronization Logic: The "CoordinatedOptimizer" includes logic to perform "all-reduce" operations on gradients of column-parallel sharded weights, ensuring correct updates in a distributed setting.
Optimizer State Sharding: Implemented functionality to partition optimizer state variables (like momentum and velocity) across multiple devices, reducing memory footprint per device.
Comprehensive Testing: New test files ("autoconfig_test.py", "coordinated_optimizer_test.py") were added to validate the automatic sharding configuration and the distributed optimizer's behavior, including handling of replicated and sharded states, and serialization.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature	Command	Description
Code Review	`/gemini review`	Performs a code review for the current pull request in its current state.
Pull Request Summary	`/gemini summary`	Provides a summary of the current pull request in its current state.
Comment	@gemini-code-assist	Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help	`/gemini help`	Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

gemini-code-assist

Code Review

This pull request introduces significant new functionality for tensor parallelism autosharding in Keras, including modules for automatic configuration and a coordinated optimizer. The implementation is well-structured, with new logic for analyzing models, generating sharding plans, and synchronizing gradients. However, I've identified a few issues that need attention. There is a critical bug in the CoordinatedOptimizer where a method for applying gradients with sharded states is called but not defined. I also found a couple of high-severity issues related to incorrect logic for matching optimizer states and gathering sharded parameters, which could lead to runtime errors or incorrect behavior. Additionally, there are some medium-severity issues regarding code clarity, such as unused parameters. The accompanying tests are a good start but do not cover the code path with the critical bug.

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

keras/src/distribution/tensor_parallel/sharding_keras.py

keras/src/distribution/tensor_parallel/autoconfig.py

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

codecov-commenter · 2025-10-15T18:35:54Z

Codecov Report

❌ Patch coverage is 59.32642% with 157 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.49%. Comparing base (a62a4a3) to head (3a4af33).
⚠️ Report is 19 commits behind head on master.

Files with missing lines	Patch %	Lines
...tribution/tensor_parallel/coordinated_optimizer.py	55.50%	87 Missing and 14 partials ⚠️
...ras/src/distribution/tensor_parallel/autoconfig.py	76.53%	11 Missing and 12 partials ⚠️
.../src/distribution/tensor_parallel/tensor_layout.py	46.15%	20 Missing and 1 partial ⚠️
keras/src/backend/jax/distributed_backend.py	29.41%	12 Missing ⚠️

Additional details and impacted files

@@            Coverage Diff             @@
##           master   #21707      +/-   ##
==========================================
- Coverage   82.59%   82.49%   -0.11%     
==========================================
  Files         572      576       +4     
  Lines       58322    58923     +601     
  Branches     9130     9236     +106     
==========================================
+ Hits        48173    48606     +433     
- Misses       7818     7965     +147     
- Partials     2331     2352      +21

Flag	Coverage Δ
keras	`82.29% <59.32%> (-0.11%)`	⬇️
keras-jax	`63.17% <58.29%> (-0.14%)`	⬇️
keras-numpy	`57.44% <36.01%> (-0.22%)`	⬇️
keras-openvino	`34.61% <36.01%> (+0.29%)`	⬆️
keras-tensorflow	`63.78% <36.01%> (-0.27%)`	⬇️
keras-torch	`63.33% <36.01%> (-0.31%)`	⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:

❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

keras/src/distribution/tensor_parallel/autoconfig.py

hertschuh · 2025-10-15T22:21:06Z

keras/src/distribution/tensor_parallel/autoconfig.py

+    if id(current_layer) in processed_layers:
+        return
+    processed_layers.add(id(current_layer))


Per my comment below about not needing a recursion, this is not needed

hertschuh · 2025-10-15T22:32:37Z

keras/src/distribution/tensor_parallel/autoconfig.py

+    processed_layers.add(id(current_layer))
+
+    name = current_layer.name
+    full_name = f"{prefix}.{name}" if prefix else name


Because you will never really recurse, the prefix won't work.

hertschuh · 2025-10-15T23:41:06Z

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

+        self._variable_to_slot_name = {}
+        opt_name = self.base_optimizer.name
+
+        normalized_params = sorted(


I have a hard time following what the code in this method is doing. I think it's try to re-pair the optimizer variables with the corresponding model variables.

I think it would be easier to capture that information in BaseOptimizer.add_variable_from_reference. Today we have _get_variable_index, but we need something more specific.

Also what do you call a slot?

hertschuh · 2025-10-16T00:03:43Z

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

+            numpy_grad = ops.convert_to_numpy(gradients[0])
+            synced_numpy = all_reduce_fn(numpy_grad, op="mean")
+            synced_tensor = ops.convert_to_tensor(synced_numpy)


numpy_grad = ops.convert_to_numpy(gradients[0])

This is going to copy the gradient tensor from GPU/TPU to CPU memory. At that point, there is no sharding anymore because it's going to gather all shards from all GPUs into one single recombined tensor on CPU.

synced_numpy = all_reduce_fn(numpy_grad, op="mean")

There is no all_reduce needed here, this is a CPU NumPy array, you might as well do np.mean.

However, what JAX will do is copy the full tensor (unsharded) from CPU to device 0 and perform the mean on that one device.

synced_tensor = ops.convert_to_tensor(synced_numpy)

synced_numpy is already a JAX array per my comment above, so this is a no-op.

So overall, I'm not sure what the intent is, but it look like this is not doing what you think it's doing. In particular, it's moving gradient back and forth to CPU, which will reduce the throughput.

hertschuh · 2025-10-16T00:06:26Z

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

+        stacked_grads = keras.ops.stack(
+            [ops.convert_to_tensor(g) for g in gradients], axis=0
+        )
+        mean_grad = ops.mean(stacked_grads, axis=0)


Why would there be more than 1 gradient in this case?

hertschuh · 2025-10-16T00:08:07Z

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

+        mean_grad = ops.mean(stacked_grads, axis=0)
+        return [mean_grad for _ in range(len(gradients))]
+
+    def get_weights(self):


Oh weird, I wonder why get_weights is missing in BaseOptimizer.

hertschuh · 2025-10-16T00:21:44Z

keras/src/distribution/tensor_parallel/coordinated_optimizer.py

+        self._initialize_sharded_states()
+
+
+class TensorParallelOptimizer(optimizers.Optimizer):


It appears that TensorParallelOptimizer is mostly a wrapper around CoordinatedOptimizer.

Any reason to have both separate?

buildwithsuhana added 6 commits October 1, 2025 15:59

adding autoconfig and coordinated_optimizer

dd3181e

Reformatting

bcae2f6

Added sharding keras

439643b

Merge branch 'keras-team:master' into Tensor_parallel_keras_2

36edcb9

Reformatting files

b7862d9

Merge branch 'Tensor_parallel_keras_2' of https://github.com/buildwit…

e8b51f7

…hsuhana/keras into Tensor_parallel_keras_2

google-ml-butler bot added the size:XL label Oct 2, 2025

google-ml-butler bot assigned gbaned Oct 2, 2025

gemini-code-assist bot reviewed Oct 2, 2025

View reviewed changes

buildwithsuhana marked this pull request as draft October 2, 2025 17:28

buildwithsuhana added 6 commits October 3, 2025 11:55

Reformatting according to changes in distributed_backend

3383dec

Reformatting according to changes in distributed_backend

5824c66

Refactoring the code

9cf5c7f

refactoring

996a154

refactoring

31994da

Testing PR1&2

8124b08

buildwithsuhana added 2 commits October 16, 2025 00:06

Testing PR1&2

3a4af33

removing pr1 contents

ec0009a

buildwithsuhana marked this pull request as ready for review October 15, 2025 20:30

hertschuh reviewed Oct 16, 2025

View reviewed changes

		self._initialize_sharded_states()


		class TensorParallelOptimizer(optimizers.Optimizer):

Add Autoconfig, Coordinated_Optimizer and Sharding keras implementations for Tensor Parallel Autosharding #21707

Are you sure you want to change the base?

Add Autoconfig, Coordinated_Optimizer and Sharding keras implementations for Tensor Parallel Autosharding #21707

Uh oh!

Conversation

buildwithsuhana commented Oct 2, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

gemini-code-assist bot commented Oct 2, 2025

Summary of Changes

Highlights

Footnotes

Uh oh!

gemini-code-assist bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

codecov-commenter commented Oct 15, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Codecov Report

Uh oh!

Uh oh!

Uh oh!

Uh oh!

hertschuh Oct 15, 2025

Choose a reason for hiding this comment

Uh oh!

hertschuh Oct 15, 2025

Choose a reason for hiding this comment

Uh oh!

hertschuh Oct 15, 2025

Choose a reason for hiding this comment

Uh oh!

hertschuh Oct 16, 2025

Choose a reason for hiding this comment

Uh oh!

hertschuh Oct 16, 2025

Choose a reason for hiding this comment

Uh oh!

hertschuh Oct 16, 2025

Choose a reason for hiding this comment

Uh oh!

hertschuh Oct 16, 2025

Choose a reason for hiding this comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

buildwithsuhana commented Oct 2, 2025 •

edited

Loading

codecov-commenter commented Oct 15, 2025 •

edited

Loading