Add VLABench (language-conditioned manipulation with long-horizon reasoning) as a new simulation benchmark, following the established LIBERO/MetaWorld patterns. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>