Platforms & Datasets
Platforms & Datasets

ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
ManiSoft provides an open vision-language manipulation platform for soft continuum robots, connecting language-conditioned policies with deformable-arm simulation and real-world evaluation.
- • Open Platform: An end-to-end simulator, benchmark, dataset, and policy stack for soft continuum manipulation
- • Expert Dataset: 6,300 expert demonstrations spanning diverse language-conditioned manipulation tasks
- • Sim-to-Real: Evaluation tools for transferring manipulation policies from simulation to physical robots

Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization
GAGeo
GAGeo unifies cross-view object geo-localization in a geometry-aware single-stage framework, supported by CMA-Loc, a large-scale benchmark with ground, drone, and satellite views.
- • Unified Framework: Single-stage localization across heterogeneous ground, drone, and satellite views
- • Geometry-Aware Reasoning: Explicit geometric cues improve cross-view localization beyond appearance matching
- • CMA-Loc Benchmark: 223,767 cross-view pairs covering 77,200 locations

Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
TopoGPT
TopoGPT formulates lane topology reasoning as autoregressive graph generation and introduces geometry priors to produce coherent lane structures in complex driving scenes.
- • Autoregressive Reasoning: Generates lane topology as a structured sequence instead of isolated detections
- • Geometry Prior: Uses scene geometry to improve topological consistency and spatial accuracy
- • Large-Scale Training: Built on 3.3 million lane-graph scenes

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance
SVCBench evaluates whether multimodal models can maintain spatial-temporal state while counting events in streaming video, where the full visual history is not available at once.
- • Streaming Evaluation: Tests online state maintenance instead of offline access to complete videos
- • Benchmark Scale: 406 videos, 1,000 questions, 4,576 query points, and 10,071 annotations
- • Counting Reasoning: Measures fine-grained event tracking across long temporal contexts

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
ViSTR-Bench measures continuous visual reasoning in dynamic scenes, focusing on whether multimodal large language models can follow evolving cues rather than rely on sparse keyframes.
- • Continuous Cues: Evaluates reasoning over evolving visual evidence in dynamic scenes
- • Diverse Tasks: 15 subtasks covering spatial, temporal, and state-transition reasoning
- • Open Benchmark: 1,340 visual question-answer pairs with public code and data

EVA-Client: A Unified Data Collection, Inference, and Deployment Framework for Embodied Policies on Real Robots
EVA-Client integrates robot data collection, policy inference, evaluation, and deployment into one open framework, reducing the engineering overhead of embodied-AI experiments on physical robots.
- • Unified Workflow: Connects data collection, inference, evaluation, and deployment
- • Real-Robot Support: Designed for reproducible embodied-policy experiments on physical platforms
- • Open Framework: Modular client architecture for extending policies and robot embodiments
GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation
GE-Sim 2.0 is an open closed-loop video world simulator for robotic manipulation, combining action-conditioned generation, scalable evaluation, and fast interactive rollout.
- • Closed-Loop Simulation: Generates action-conditioned visual feedback for interactive policy evaluation
- • WorldArena Evaluation: A unified benchmark for comparing video world simulators
- • Open Ecosystem: Public code and model weights support reproducible research and deployment

OctoNav: Towards Generalist Embodied Navigation
We introduce OctoNav, a generalist embodied navigation model designed to handle diverse navigation tasks across different environments. By leveraging large-scale pre-training and multi-modal integration, OctoNav achieves robust performance in complex indoor and outdoor scenarios.
- • Generalist Navigation: Unified model architecture capable of handling multiple embodied navigation tasks
- • Multi-modal Integration: Effective fusion of visual perception and linguistic instructions for robust control
- • Cross-Environment Performance: Demonstrates strong generalization across diverse and unseen navigation scenarios

Towards Realistic UAV Vision-Language Navigation: Platform, Benchmark, and Methodology
We present a comprehensive platform for UAV vision-language navigation that bridges the gap between simulation and real-world deployment. Our benchmark provides realistic scenarios and evaluation metrics for advancing autonomous UAV systems.
- • Platform: Comprehensive simulation and real-world testing environment for UAV navigation
- • Benchmark: Realistic scenarios with standardized evaluation metrics
- • Methodology: Novel approach bridging sim-to-real gap in vision-language navigation

VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
A large-scale dataset designed for fine-grained video reasoning through chain-of-thought methodology. We introduce core frame selection techniques to enhance video understanding and reasoning capabilities.
- • Dataset Scale: Large-scale collection with diverse video scenarios and reasoning tasks
- • Chain-of-Thought: Structured reasoning methodology for fine-grained video understanding
- • Core Frame Selection: Novel technique to identify and leverage key frames for efficient reasoning

LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
A multimodal large language model specifically designed for fine-grained spatial-temporal understanding. LLaVA-ST enables precise comprehension of complex spatial and temporal relationships in visual content.
- • Multimodal Architecture: Advanced integration of vision and language for spatial-temporal reasoning
- • Fine-Grained Understanding: Precise comprehension of complex spatial and temporal relationships
- • Applications: Enhanced performance in video understanding and visual question answering

Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localization
We transform drone videos into bird's-eye-view representations for improved geo-localization. This approach enables accurate spatial understanding and localization from aerial video footage.
- • BEV Transformation: Novel method to convert drone videos into bird's-eye-view representations
- • Geo-localization: Accurate spatial understanding and localization from aerial footage
- • Video-based Approach: Leverages temporal information for robust localization performance

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
An innovative dual-system approach for UAV-based vision and language navigation. AeroDuo combines visual perception with natural language understanding to enable intuitive UAV control and navigation.
- • Dual-System Design: Innovative architecture combining visual and linguistic processing
- • Vision-Language Integration: Seamless fusion of visual perception and natural language understanding
- • Intuitive Control: Natural language commands for user-friendly UAV navigation

RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
A comprehensive benchmark for evaluating long-horizon robotic manipulation tasks. RoboCerebra provides standardized evaluation protocols and diverse scenarios to advance research in embodied AI and robotics.
- • Large-scale Benchmark: Comprehensive evaluation framework for long-horizon robotic manipulation
- • Standardized Protocols: Consistent evaluation metrics and methodologies across diverse scenarios
- • Research Advancement: Advances embodied AI and robotics research through systematic evaluation
Hi AirStar: Guide Me to the Badminton Court
An interactive demonstration system that guides users to specific locations using natural language commands. This demo showcases practical applications of vision-language navigation in real-world scenarios.
- • Interactive System: Real-time guidance through natural language interaction
- • Practical Application: Demonstrates vision-language navigation in real-world settings
- • User Experience: Intuitive location guidance with conversational interface