3A-VLA: Abstraction-Aligned Action Learning for Vision-Language Agents in 3D Game Worlds

Anonymous Submission

Code Data
3A-VLA Framework Overview

Overview. 3A-VLA learns abstraction-aligned actions by coupling Intention Abstraction (IA) and Environment Semantics Abstraction (ESA) with Abstraction Alignment Reweighting (AAR). During training, IA and ESA expose sparse task primitives from language and vision, while AAR reweights action supervision according to intention–environment alignment. At inference, the induced sparsity enables low-latency control by keeping only top-k task-critical visual tokens.

Abstract

Despite advances in Visual-Language-Action (VLA) models, they remain limited in highly dynamic game settings—such as 3D open worlds and competitive PvP—where agents must efficiently extract sparse, actionable signals from dense visual input while integrating multi-modal cues to track off-screen high-value targets in real time. To tackle this, we introduce 3A-VLA (Abstraction-Aligned Action VLA), a framework that grounds action learning in explicit intention and environment abstractions rather than superficial pattern matching. We introduce dual task-agnostic abstractions: Intention Abstraction (IA), which condenses verbose instructions and reasoning into explicit semantic primitives; Environment Semantics Abstraction (ESA), which structures dense visual streams into a spatial-functional affordance representation to guide grounded action. We further propose an Abstraction Alignment Reweighting (AAR) module that adaptively reweights the action imitation loss. It uses a continuous intention-environment alignment signal to emphasize reliable action supervision when abstractions agree and reduce the influence of ambiguous demonstrations when they diverge, thereby learning a policy that balances fine-grained control and high-level reasoning without manual rules. Extensive experiments show that 3A-VLA establishes new state-of-the-art results in both open-world (Minecraft) and competitive PvP (Game for Peace) settings. It also demonstrates strong zero-shot generalization to high-fidelity games across different domains, including Valorant, CS2, GTA V, Elden Ring, and Mount & Blade.

Benchmark on Game for Peace

We establish a taxonomy of six atomic tasks that encapsulate the complete lifecycle of a battle royale match at an intermediate difficulty level (Gold and Silver tiers).

Precision Parachuting

Controlling descent trajectory to land within a minimal radius of a designated waypoint.

Resource Scavenging

Exploring the environment to identify, pick up, and collect essential loot such as weapons and armor.

Combat Engagement

Detecting adversaries, tracking their movement, and compensating for recoil to deliver lethal damage in encounters.

Teammate Revival

Identifying knocked-down teammates, navigating to their location, and reviving them to restore their combat status.

Vehicle Acquisition

Searching for, locating, and boarding available vehicles to secure strategic mobility across the battlefield.

Strategic Rotation

Navigating towards the shrinking safe zone while avoiding obstacles under strict time constraints.

Zero-shot Generalization to Other Games

Minecraft

Valorant

Grand Theft Auto V

Elden Ring

Mount & Blade