feat: Major optimization and robustness improvements (#31)
This release introduces significant performance optimizations, memory efficiency improvements, and enhanced system robustness: 🚀 Performance Optimizations: - Add three new fused CUDA kernels (rope_index, rot_pos_emb, get_window_index) for accelerated multimodal preprocessing - Implement FSDP2 support for distributed training with improved memory efficiency - Add Torch.compile integration for additional performance gains - Optimize memory usage: reduce peak allocation from 48GB to 24GB on 8-GPU setup 🔧 System Robustness: - Fix missing token position inputs in prediction pipeline - Add type-robust negation operations in RoPE CUDA kernels (half/bfloat16 support) - Fix dataset root parameter initialization in LeRobot data loader - Enhanced error handling and input validation across fusion operators 📚 Documentation & Usability: - Add comprehensive memory usage benchmarks and hardware recommendations - Update citation format with proper arXiv reference - Improve training configuration documentation with quick start guide - Add detailed API documentation for new fusion operators 🛠️ Technical Details: - Version bump to 1.0.1 - New CUDA kernels: rope_index.cu, rot_pos.cu, window_index.cu - FSDP2 state dict loading with distribute_tensor support - Enhanced multimodal RoPE with 3D position encoding - Window attention optimization for Vision Transformers Breaking Changes: None - all changes are backward compatible
This commit is contained in:
@@ -36,6 +36,9 @@ ext_modules = [
|
||||
"csrc/dual_asym_grouped_gemm.cu",
|
||||
"csrc/permute.cu",
|
||||
"csrc/rope.cu",
|
||||
"csrc/rope_index.cu",
|
||||
"csrc/rot_pos.cu",
|
||||
"csrc/window_index.cu",
|
||||
],
|
||||
include_dirs=[f"{cwd}/3rdparty/cutlass/include/", f"{cwd}/csrc"],
|
||||
extra_compile_args={
|
||||
@@ -47,7 +50,7 @@ ext_modules = [
|
||||
|
||||
setup(
|
||||
name="wall_x",
|
||||
version="1.0.0",
|
||||
version="1.0.1",
|
||||
author="X2Robot Team",
|
||||
classifiers=[
|
||||
"Programming Language :: Python :: 3",
|
||||
|
||||
Reference in New Issue
Block a user