Camera localization methods are mostly multi-stage pipelines relying on dedicated and complex maps that are costly to construct, store, and maintain over time. Alternatively, well-structured maps that organize scenes with compact yet expressive primitives are often readily available, providing rich structural and semantic cues that humans can interpret. However, their potential for direct camera localization remains largely unexplored due to the significant modality gap with visual observations. This motivates a new concept for localization, which estimates camera poses by grounding query observations to map primitives. In this paper, we instantiate the concept in the context of planar primitives with PGGT, the Primitive-Grounded Geometry Transformer network trained to predict multiple localization targets from a single query image and a set of planar map primitives, enabling feed-forward 6-DoF camera localization that generalizes to unseen environments. Based on this flexible and scalable architecture, our method pushes the state-of-the-art localization performance on both detailed planar surface maps and minimal architectural layouts. The code and models will be made publicly available.