AdithyaSK HF Staff commited on
Commit
c8d297c
·
verified ·
1 Parent(s): 04379e8

Remove committed setuptools build/ and egg-info; they shadow the real sources when the Space is pip-installed

Browse files
Files changed (28) hide show
  1. env/openenv_geoguesser_env.egg-info/PKG-INFO +0 -233
  2. env/openenv_geoguesser_env.egg-info/SOURCES.txt +0 -21
  3. env/openenv_geoguesser_env.egg-info/dependency_links.txt +0 -1
  4. env/openenv_geoguesser_env.egg-info/entry_points.txt +0 -2
  5. env/openenv_geoguesser_env.egg-info/requires.txt +0 -14
  6. env/openenv_geoguesser_env.egg-info/top_level.txt +0 -1
  7. geoguesser_env/build/lib/geoguesser_env/__init__.py +0 -49
  8. geoguesser_env/build/lib/geoguesser_env/client.py +0 -143
  9. geoguesser_env/build/lib/geoguesser_env/harness.py +0 -440
  10. geoguesser_env/build/lib/geoguesser_env/models.py +0 -462
  11. geoguesser_env/build/lib/geoguesser_env/server/__init__.py +0 -3
  12. geoguesser_env/build/lib/geoguesser_env/server/app.py +0 -349
  13. geoguesser_env/build/lib/geoguesser_env/server/backends/__init__.py +0 -7
  14. geoguesser_env/build/lib/geoguesser_env/server/backends/base.py +0 -127
  15. geoguesser_env/build/lib/geoguesser_env/server/backends/panorama.py +0 -399
  16. geoguesser_env/build/lib/geoguesser_env/server/geoguesser_environment.py +0 -998
  17. geoguesser_env/build/lib/geoguesser_env/server/gradio_ui.py +0 -1044
  18. geoguesser_env/build/lib/geoguesser_env/server/parser.py +0 -149
  19. geoguesser_env/build/lib/geoguesser_env/server/render/__init__.py +0 -3
  20. geoguesser_env/build/lib/geoguesser_env/server/render/minimap.py +0 -824
  21. geoguesser_env/build/lib/geoguesser_env/server/render/pano.py +0 -108
  22. geoguesser_env/build/lib/geoguesser_env/server/scoring.py +0 -233
  23. geoguesser_env/openenv_geoguesser_env.egg-info/PKG-INFO +0 -760
  24. geoguesser_env/openenv_geoguesser_env.egg-info/SOURCES.txt +0 -29
  25. geoguesser_env/openenv_geoguesser_env.egg-info/dependency_links.txt +0 -1
  26. geoguesser_env/openenv_geoguesser_env.egg-info/entry_points.txt +0 -2
  27. geoguesser_env/openenv_geoguesser_env.egg-info/requires.txt +0 -14
  28. geoguesser_env/openenv_geoguesser_env.egg-info/top_level.txt +0 -1
env/openenv_geoguesser_env.egg-info/PKG-INFO DELETED
@@ -1,233 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: openenv-geoguesser-env
3
- Version: 0.1.0
4
- Summary: GeoGuessr-style visual geolocation environment for OpenEnv
5
- Requires-Python: >=3.10
6
- Description-Content-Type: text/markdown
7
- Requires-Dist: openenv>=0.3.1
8
- Requires-Dist: fastapi>=0.115.0
9
- Requires-Dist: pydantic>=2.0.0
10
- Requires-Dist: uvicorn>=0.24.0
11
- Requires-Dist: fastmcp>=2.0.0
12
- Requires-Dist: pillow>=10.0.0
13
- Requires-Dist: numpy>=1.24.0
14
- Requires-Dist: matplotlib>=3.7.0
15
- Provides-Extra: ui
16
- Requires-Dist: gradio>=4.0.0; extra == "ui"
17
- Provides-Extra: dev
18
- Requires-Dist: pytest>=8.0.0; extra == "dev"
19
-
20
- # GeoGuesser
21
-
22
- A GeoGuessr-style visual geolocation environment. The agent is dropped at an
23
- unknown street-level location, looks around, walks along the road, pins
24
- candidate coordinates on a map to check itself, and commits to a final guess.
25
- Reward is distance-based, using the game's own scoring curve.
26
-
27
- Independent open-source project, unaffiliated with GeoGuessr AB. Imagery comes
28
- from Mapillary contributors under CC-BY-SA-4.0.
29
-
30
- ## Quick start
31
-
32
- ```bash
33
- # 1. Build a task index (needs a free Mapillary token with READ scope)
34
- export MAPILLARY_API_KEY="MLY|..."
35
- cd envs/geoguesser_env
36
- uv run python scripts/build_pano_tasks.py --tasks 100 --frames 24
37
-
38
- # 2. Run the server
39
- uv run --project . server # http://localhost:8000
40
- ```
41
-
42
- ```python
43
- from geoguesser_env import GeoGuesserEnv, GuessAction, LookAction, PinAction
44
-
45
- env = GeoGuesserEnv(base_url="http://localhost:8000")
46
-
47
- result = env.reset(task_index=7) # byte-identical on repeat
48
- print(result.observation.prompt)
49
-
50
- result = env.step(LookAction(heading_deg=90, fov_deg=45))
51
- result = env.step(PinAction(lat=-16.5, lon=-68.1))
52
- print(result.observation.feedback)
53
- # Pin 1 placed at -16.5000, -68.1000 - Bolivia (South America).
54
- # Nearest major city: La Paz, ~5 km E. 10 actions left.
55
-
56
- result = env.step(GuessAction(response="Altiplano. <guess>-16.49, -68.12</guess>"))
57
- print(result.reward, result.observation.distance_km)
58
- ```
59
-
60
- ## Tools
61
-
62
- | Tool | What it does | Cost |
63
- |------|--------------|------|
64
- | `look(heading_deg, pitch_deg, fov_deg)` | Render a view. Heading is absolute, `0` is true north | −0.01 |
65
- | `pan(delta_deg)` | Turn relative to the current heading | −0.01 |
66
- | `zoom(fov_deg)` | Narrow the field of view; around 30 reads distant signs | −0.01 |
67
- | `move(direction, meters)` | Walk the captured road; reports distance actually travelled | −0.05 |
68
- | `place_pin(lat, lon, label)` | Pin a candidate and see where it falls on the map | −0.02 |
69
- | `view_map(lat, lon, span_deg)` | Pan and zoom the map without pinning | −0.01 |
70
- | `list_pins()` / `clear_pins()` | Review or drop candidates | free |
71
- | `measure(lat_a, lon_a, lat_b, lon_b)` | Distance between two of your own points | free |
72
- | `reverse_geocode(lat, lon)` | Name the country and nearest city at a coordinate | free |
73
- | `submit_guess(lat, lon, ...)` | Commit the answer. Terminal | — |
74
-
75
- Tools the backend cannot serve are **not registered**, so the agent never sees
76
- a tool that always fails.
77
-
78
- ### Pinning tells you where you pointed, not whether you are right
79
-
80
- `place_pin` returns a rendered map and a description of the pinned location:
81
- country, subregion, nearest city with distance and bearing, and the distance
82
- to the agent's own earlier pins. It reveals nothing about the target.
83
-
84
- That restraint is deliberate. Any signal about the truth — a distance, a
85
- warmer/colder hint — would make binary search the optimal policy, and the
86
- environment would measure bisection rather than geographic reasoning. Distance
87
- and score arrive only from `submit_guess`.
88
-
89
- ## Reward
90
-
91
- ```
92
- geo = exp(-distance_km / 1492.7) # GeoGuessr's curve, in [0, 1]
93
- partial = 0.15 * country_hit + 0.10 * region_hit # when hierarchical
94
- cost = 0.01*looks + 0.01*maps + 0.02*pins + 0.05*moves
95
- reward = clip(geo + partial, 0, 1) - cost
96
- ```
97
-
98
- An unparseable or out-of-range guess scores `0.0` and says why. Parsing
99
- accepts what models actually emit: decimal pairs, DMS (`48°51'29"N`), labelled
100
- `lat:`/`lon:`, JSON, and `<guess>` tags.
101
-
102
- ## Reproducibility
103
-
104
- ```python
105
- env.reset(task_index=7) # exact task, byte-identical observation -> GRPO, eval
106
- env.reset(seed=42) # tasks[42 % n_tasks] -> replay
107
- env.reset() # random task, index in metadata -> UI
108
- ```
109
-
110
- Byte-identical repeats hold because panorama bytes come from a local cache
111
- rather than an expiring CDN URL, reprojection is pure numpy with integer
112
- sampling, and the initial heading is pinned to each panorama's own
113
- `compass_angle`.
114
-
115
- An eval score is only meaningful alongside its provenance — the env version,
116
- the task index, and `GEODATA_VERSION` from `server/render/minimap.py`, since
117
- the bundled vectors determine the reverse-geocode text the agent sees.
118
-
119
- ## Training and collection
120
-
121
- The environment plugs into `openenv.core.harness`, so a rollout function and a
122
- collector come for free:
123
-
124
- ```python
125
- from geoguesser_env import GeoGuesserEnv
126
- from geoguesser_env.harness import GeoGuesserSessionFactory, load_tasks
127
-
128
- tasks = load_tasks("tasks/pano_v1.jsonl", repeat=16) # GRPO group of 16
129
- factory = GeoGuesserSessionFactory(
130
- lambda: GeoGuesserEnv(base_url="http://localhost:8000")
131
- )
132
- ```
133
-
134
- See `examples/geoguesser_rollout.py` for a scripted rollout and
135
- `examples/geoguesser_collect.py` for JSONL collection with resume.
136
-
137
- ## Human play
138
-
139
- A five-round game, 5,000 points a round on the same curve the environment
140
- rewards, so a human score is directly comparable to GeoGuessr intuition and to
141
- the agent's reward (both are shown).
142
-
143
- ```bash
144
- uv run --project . server
145
- # then open http://localhost:8000/geoguesser/play
146
- ```
147
-
148
- The page stands alone at `/geoguesser/play` and is also embedded in the Gradio
149
- playground's **Custom** tab when the web interface is enabled:
150
-
151
- ```bash
152
- ENABLE_WEB_INTERFACE=true uv run --project . server # http://localhost:8000/web/
153
- ```
154
-
155
- Drag to look around, scroll to zoom, **M** toggles a larger map, **Enter**
156
- submits and then advances. On submit the map takes the screen and draws the
157
- line between guess and truth, exactly like the game; the result bar shows
158
- distance, points, env reward and the true location, and a scoreboard breaks
159
- down all five rounds at the end.
160
-
161
- Panoramas are rendered by Pannellum and the map by MapLibre over OpenFreeMap
162
- tiles — no API key, no request limits. The imagery credit line names the
163
- Mapillary contributor, which the CC-BY-SA licence requires.
164
-
165
- The page has to be a standalone document rather than a Gradio `gr.HTML`
166
- fragment: `gr.HTML` inserts markup without executing `<script>` tags, so the
167
- viewers never initialise and the panel renders blank with no error anywhere.
168
-
169
- Extra routes, all local:
170
-
171
- | Route | Returns |
172
- |-------|---------|
173
- | `/geoguesser/play` | the play page |
174
- | `/geoguesser/tasks` | `{"n_tasks": N}` |
175
- | `/geoguesser/task/{i}` | task metadata, including ground truth for the human UI |
176
- | `/geoguesser/pano/{i}` | the raw equirectangular panorama |
177
-
178
- The human map uses live tiles; the agent's map stays the offline Natural Earth
179
- render, so the agent keeps a determinism the browser does not need. Note that
180
- `/geoguesser/task/{i}` exposes ground truth — it exists for a person playing in
181
- their own browser, and agent observations still withhold it until the guess.
182
-
183
- ## Configuration
184
-
185
- | Variable | Default | Meaning |
186
- |----------|---------|---------|
187
- | `GEOGUESSER_INDEX` | `tasks/pano_v1.jsonl` | Task index to load |
188
- | `GEOGUESSER_CACHE` | `data/panos` | Panorama cache directory |
189
- | `GEOGUESSER_EPISODE_MODE` | `agentic` | `agentic`, `single_shot` or `nmpz` |
190
- | `GEOGUESSER_MAX_STEPS` | `12` | Actions before the episode is cut off |
191
- | `GEOGUESSER_REWARD_MODE` | `coords` | `coords` or `country_only` |
192
- | `GEOGUESSER_HIERARCHICAL` | `0` | Add country and region partial credit |
193
- | `GEOGUESSER_VIEW_SIZE` | `640` | Edge length of rendered views |
194
- | `GEOGUESSER_ALLOW_FETCH` | `1` | Whether a cache miss may reach the API |
195
- | `MAPILLARY_API_KEY` | — | Needed by the builder, and only on a cache miss |
196
-
197
- ## Data
198
-
199
- The committed index holds **100 tasks across 47 countries and all six
200
- continents**, one per Mapillary sequence, captured between 2017 and 2026:
201
-
202
- | | |
203
- |---|---|
204
- | Tasks | 100 (`task_index` 0-99) |
205
- | Countries | 47, capped at 4 tasks each |
206
- | Continents | Europe 34, Asia 27, South America 23, North America 11, Oceania 3, Africa 2 |
207
- | Frames per task | 23.2 mean (3 min, 24 max), ~3.3 m apart |
208
- | Index size | 417 KB, committed |
209
- | Cache | 37 MB for the 100 start frames; ~0.6 GB fully warmed |
210
-
211
- The index is self-contained: every frame's coordinates, heading and capture
212
- date live in `tasks/pano_v1.jsonl`, so the movement graph resolves offline.
213
- Only image bytes are fetched, and only on a cache miss, because Mapillary
214
- `thumb_*_url` values are expiring signed URLs that cannot be stored.
215
-
216
- Coverage is uneven and worth knowing about. Probing 45 Street-View
217
- coordinates found any Mapillary imagery at 21 and a 360-degree panorama at
218
- only 7, heavily clustered. Panorama-first discovery is therefore the only
219
- approach that works — roughly 5% of probe points yield a usable sequence, so
220
- reaching 100 tasks took two passes with different seeds, merged by
221
- `scripts/merge_task_indexes.py`. Africa and Oceania are thin because 360-degree
222
- contributors are; that is a property of the source, documented rather than
223
- papered over.
224
-
225
- ## Known gaps versus the real game
226
-
227
- Movement follows captured sequences and stops where one ends. There is no
228
- multi-round cumulative score, no wall-clock timer (a step budget stands in for
229
- it), and no satellite layer on the guess map. Coverage hints and web search are
230
- deliberately excluded: the first is a crutch, the second turns the task into
231
- retrieval.
232
-
233
- See [DESIGN.md](DESIGN.md) for the reasoning behind these choices.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
env/openenv_geoguesser_env.egg-info/SOURCES.txt DELETED
@@ -1,21 +0,0 @@
1
- README.md
2
- pyproject.toml
3
- openenv_geoguesser_env.egg-info/PKG-INFO
4
- openenv_geoguesser_env.egg-info/SOURCES.txt
5
- openenv_geoguesser_env.egg-info/dependency_links.txt
6
- openenv_geoguesser_env.egg-info/entry_points.txt
7
- openenv_geoguesser_env.egg-info/requires.txt
8
- openenv_geoguesser_env.egg-info/top_level.txt
9
- server/__init__.py
10
- server/app.py
11
- server/geoguesser_environment.py
12
- server/gradio_ui.py
13
- server/parser.py
14
- server/scoring.py
15
- server/backends/__init__.py
16
- server/backends/base.py
17
- server/backends/panorama.py
18
- server/render/__init__.py
19
- server/render/minimap.py
20
- server/render/pano.py
21
- tests/test_geoguesser_env.py
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
env/openenv_geoguesser_env.egg-info/dependency_links.txt DELETED
@@ -1 +0,0 @@
1
-
 
 
env/openenv_geoguesser_env.egg-info/entry_points.txt DELETED
@@ -1,2 +0,0 @@
1
- [console_scripts]
2
- server = server.app:main
 
 
 
env/openenv_geoguesser_env.egg-info/requires.txt DELETED
@@ -1,14 +0,0 @@
1
- openenv>=0.3.1
2
- fastapi>=0.115.0
3
- pydantic>=2.0.0
4
- uvicorn>=0.24.0
5
- fastmcp>=2.0.0
6
- pillow>=10.0.0
7
- numpy>=1.24.0
8
- matplotlib>=3.7.0
9
-
10
- [dev]
11
- pytest>=8.0.0
12
-
13
- [ui]
14
- gradio>=4.0.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
env/openenv_geoguesser_env.egg-info/top_level.txt DELETED
@@ -1 +0,0 @@
1
- server
 
 
geoguesser_env/build/lib/geoguesser_env/__init__.py DELETED
@@ -1,49 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """GeoGuesser: a GeoGuessr-style visual geolocation environment.
4
-
5
- Independent open-source project, unaffiliated with GeoGuessr AB. Imagery comes
6
- from Mapillary contributors under CC-BY-SA-4.0.
7
- """
8
-
9
- from .client import GeoGuesserEnv
10
- from .models import (
11
- EpisodeMode,
12
- from_wire,
13
- GeoGuesserAction,
14
- GeoGuesserObservation,
15
- GeoGuesserState,
16
- GuessAction,
17
- LookAction,
18
- MeasureAction,
19
- MoveAction,
20
- PanAction,
21
- Pin,
22
- PinAction,
23
- RewardMode,
24
- to_wire,
25
- TypedAction,
26
- ViewMapAction,
27
- ZoomAction,
28
- )
29
-
30
- __all__ = [
31
- "EpisodeMode",
32
- "GeoGuesserAction",
33
- "GeoGuesserEnv",
34
- "GeoGuesserObservation",
35
- "GeoGuesserState",
36
- "GuessAction",
37
- "LookAction",
38
- "MeasureAction",
39
- "MoveAction",
40
- "PanAction",
41
- "Pin",
42
- "PinAction",
43
- "RewardMode",
44
- "TypedAction",
45
- "ViewMapAction",
46
- "ZoomAction",
47
- "from_wire",
48
- "to_wire",
49
- ]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/client.py DELETED
@@ -1,143 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """HTTP client for the GeoGuesser environment."""
4
-
5
- from __future__ import annotations
6
-
7
- from typing import Any, Dict
8
-
9
- from openenv.core.client_types import StepResult
10
- from openenv.core.env_client import EnvClient
11
-
12
- from .models import (
13
- GeoGuesserAction,
14
- GeoGuesserObservation,
15
- GeoGuesserState,
16
- to_wire,
17
- TypedAction,
18
- )
19
-
20
-
21
- ENV_NAME = "geoguesser_env"
22
- """Task API routes are namespaced by the `env_name` the server registers."""
23
-
24
-
25
- class GeoGuesserEnv(
26
- EnvClient[GeoGuesserAction, GeoGuesserObservation, GeoGuesserState]
27
- ):
28
- """
29
- Client for a running GeoGuesser environment server.
30
-
31
- Examples:
32
-
33
- ```python
34
- env = GeoGuesserEnv(base_url="http://localhost:8000")
35
- result = env.reset(task_index=0)
36
- result = env.step(LookAction(heading_deg=90))
37
- result = env.step(GuessAction(response("<guess>55.67, 12.57</guess>")))
38
- print(result.reward, result.observation.distance_km)
39
- ```
40
- """
41
-
42
- def _step_payload(self, action: GeoGuesserAction | TypedAction) -> Dict[str, Any]:
43
- """
44
- Serialise an action for the `/step` endpoint.
45
-
46
- Typed actions are flattened to the single wire schema the server
47
- declares; a wire action passes through unchanged.
48
- """
49
- wire = action if isinstance(action, GeoGuesserAction) else to_wire(action)
50
- return wire.model_dump(exclude_none=True)
51
-
52
- def _parse_result(self, response: Dict[str, Any]) -> StepResult:
53
- """Build a [`StepResult`] from a `/step` or `/reset` response."""
54
- observation = GeoGuesserObservation(**response["observation"])
55
- return StepResult(
56
- observation=observation,
57
- reward=response.get("reward"),
58
- done=response.get("done", False),
59
- )
60
-
61
- def _parse_state(self, response: Dict[str, Any]) -> GeoGuesserState:
62
- """Build a [`GeoGuesserState`] from a `/state` response."""
63
- return GeoGuesserState(**response)
64
-
65
- def reset(
66
- self,
67
- task_index: int | None = None,
68
- split: str | None = None,
69
- index: int | None = None,
70
- **kwargs: Any,
71
- ) -> StepResult:
72
- """
73
- Start an episode.
74
-
75
- Args:
76
- task_index (`int`, *optional*):
77
- Deprecated alias for `index`, kept so existing callers and
78
- recorded trajectories keep working.
79
- split (`str`, *optional*):
80
- Which split to draw from, as named by [`list_splits`]. Defaults
81
- to the server's default split.
82
- index (`int`, *optional*):
83
- Exact task to play within `split`. Repeated calls with the same
84
- split and index yield byte-identical observations, which is what
85
- a GRPO group needs. Omit it, and `seed`, for a random task.
86
-
87
- Returns:
88
- [`StepResult`]: The opening observation.
89
- """
90
- if index is None:
91
- index = task_index
92
- if split is not None:
93
- kwargs["split"] = split
94
- if index is not None:
95
- kwargs["index"] = index
96
- return super().reset(**kwargs)
97
-
98
- def list_splits(self) -> list[dict[str, Any]]:
99
- """
100
- Which splits the server offers, and how many tasks each holds.
101
-
102
- Core exposes the Task API over HTTP but ships no client for it, so this
103
- posts to the routes directly.
104
-
105
- Returns:
106
- `list[dict]`: Split descriptors with `name`, `type`, `num_tasks` and
107
- `default`.
108
- """
109
- return self._task_api("splits", method="GET")
110
-
111
- def num_tasks(self, split: str) -> int:
112
- """How many tasks a split holds."""
113
- return int(self._task_api("num_tasks", {"split": split})["num_tasks"])
114
-
115
- def get_task(self, split: str, index: int) -> dict[str, Any]:
116
- """
117
- Describe one task without starting an episode.
118
-
119
- The spec carries no coordinates and no country: it is metadata for
120
- whatever orchestrates a run, not a label source.
121
- """
122
- return self._task_api("task", {"split": split, "index": index})["task"]
123
-
124
- def _task_api(
125
- self,
126
- route: str,
127
- payload: dict[str, Any] | None = None,
128
- method: str = "POST",
129
- ) -> Any:
130
- """Call one core Task API route on this environment."""
131
- import json
132
- import urllib.request
133
-
134
- url = f"{self.base_url.rstrip('/')}/{ENV_NAME}/{route}"
135
- data = None if payload is None else json.dumps(payload).encode()
136
- request = urllib.request.Request(
137
- url,
138
- data=data,
139
- method=method,
140
- headers={"Content-Type": "application/json"},
141
- )
142
- with urllib.request.urlopen(request, timeout=30) as response:
143
- return json.loads(response.read())
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/harness.py DELETED
@@ -1,440 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Harness-oriented GeoGuesser session adapters.
4
-
5
- Follows the pattern in `reasoning_gym_env.harness`: expose a `GeoGuesserEnv`
6
- client as a `ResourceSession` driven through MCP-style tools, so it plugs into
7
- `openenv.core.harness` unchanged.
8
-
9
- That single adapter is what makes the rest work without extra code:
10
-
11
- - `CollectRunner(tasks=...)` walks a task list into a JSONL rollout dataset,
12
- with resume, recording which task produced each episode.
13
- - `build_harness_rollout_func(...)` yields a TRL-compatible rollout function
14
- where each prompt *is* a task — the GRPO path.
15
- - `EvalConfig` / `EvalResult` carry the provenance an eval score needs.
16
-
17
- A task is a plain dict, so it survives serialisation into an `EpisodeRecord`:
18
-
19
- ```python
20
- {"task_index": 7, "task_id": "mly-0007"}
21
- ```
22
- """
23
-
24
- from __future__ import annotations
25
-
26
- import json
27
- import pathlib
28
- from typing import Any, Callable, Iterator
29
-
30
- from openenv.core.env_server.mcp_types import Tool
31
- from openenv.core.harness import (
32
- ResourceSessionFactory,
33
- StepEnvSessionAdapter,
34
- ToolResult,
35
- VerifyResult,
36
- )
37
-
38
- from .client import GeoGuesserEnv
39
- from .models import (
40
- GuessAction,
41
- LookAction,
42
- MeasureAction,
43
- MoveAction,
44
- PanAction,
45
- PinAction,
46
- to_wire,
47
- ViewMapAction,
48
- ZoomAction,
49
- )
50
-
51
-
52
- def _number(description: str) -> dict[str, Any]:
53
- return {"type": "number", "description": description}
54
-
55
-
56
- GEOGUESSER_TOOLS: list[Tool] = [
57
- Tool(
58
- name="look",
59
- description=(
60
- "Look in a direction from where you stand. heading_deg is "
61
- "absolute, 0 being true north. Smaller fov_deg zooms in."
62
- ),
63
- input_schema={
64
- "type": "object",
65
- "properties": {
66
- "heading_deg": _number("Compass heading in degrees."),
67
- "pitch_deg": _number("Vertical angle; positive looks up."),
68
- "fov_deg": _number("Field of view in degrees, 10 to 120."),
69
- },
70
- "required": ["heading_deg"],
71
- },
72
- ),
73
- Tool(
74
- name="pan",
75
- description="Turn relative to your current heading; positive turns right.",
76
- input_schema={
77
- "type": "object",
78
- "properties": {"delta_deg": _number("Degrees to turn.")},
79
- "required": ["delta_deg"],
80
- },
81
- ),
82
- Tool(
83
- name="zoom",
84
- description="Change field of view without turning. Around 30 reads distant signs.",
85
- input_schema={
86
- "type": "object",
87
- "properties": {"fov_deg": _number("New field of view in degrees.")},
88
- "required": ["fov_deg"],
89
- },
90
- ),
91
- Tool(
92
- name="move",
93
- description=(
94
- "Walk along the captured road. Reports how far you actually "
95
- "travelled, since frame spacing is irregular."
96
- ),
97
- input_schema={
98
- "type": "object",
99
- "properties": {
100
- "direction": {
101
- "type": "string",
102
- "enum": ["forward", "backward"],
103
- "description": "Which way to walk.",
104
- },
105
- "meters": _number("Requested distance in metres."),
106
- },
107
- "required": ["direction"],
108
- },
109
- ),
110
- Tool(
111
- name="place_pin",
112
- description=(
113
- "Pin a candidate coordinate and see where it falls on the map. "
114
- "Tells you what is at that coordinate. It says nothing about "
115
- "whether you are right."
116
- ),
117
- input_schema={
118
- "type": "object",
119
- "properties": {
120
- "lat": _number("Latitude of the candidate."),
121
- "lon": _number("Longitude of the candidate."),
122
- "label": {"type": "string", "description": "Optional note."},
123
- "span_deg": _number(
124
- "Half-width of the returned map in degrees. Below about 4 "
125
- "the map adds roads, urban areas and town names."
126
- ),
127
- },
128
- "required": ["lat", "lon"],
129
- },
130
- ),
131
- Tool(
132
- name="view_map",
133
- description="Pan and zoom the map without placing a pin.",
134
- input_schema={
135
- "type": "object",
136
- "properties": {
137
- "lat": _number("Latitude at the centre of the view."),
138
- "lon": _number("Longitude at the centre of the view."),
139
- "span_deg": _number("Half-width of the window in degrees."),
140
- },
141
- "required": ["lat", "lon"],
142
- },
143
- ),
144
- Tool(
145
- name="measure",
146
- description="Distance in km between two coordinates of your own choosing. Free.",
147
- input_schema={
148
- "type": "object",
149
- "properties": {
150
- "lat_a": _number("Latitude of the first point."),
151
- "lon_a": _number("Longitude of the first point."),
152
- "lat_b": _number("Latitude of the second point."),
153
- "lon_b": _number("Longitude of the second point."),
154
- },
155
- "required": ["lat_a", "lon_a", "lat_b", "lon_b"],
156
- },
157
- ),
158
- Tool(
159
- name="submit_guess",
160
- description="Commit your final answer. This ends the episode.",
161
- input_schema={
162
- "type": "object",
163
- "properties": {
164
- "lat": _number("Latitude of your guess."),
165
- "lon": _number("Longitude of your guess."),
166
- "country": {
167
- "type": "string",
168
- "description": "Optional ISO-3166 alpha-2 code or country name.",
169
- },
170
- "confidence": _number("Optional confidence in [0, 1]."),
171
- "reasoning": {
172
- "type": "string",
173
- "description": "Optional rationale, recorded but not scored.",
174
- },
175
- },
176
- "required": ["lat", "lon"],
177
- },
178
- ),
179
- ]
180
-
181
- _ACTION_BY_TOOL: dict[str, Callable[[dict[str, Any]], Any]] = {
182
- "look": lambda a: LookAction(
183
- heading_deg=float(a["heading_deg"]),
184
- pitch_deg=float(a.get("pitch_deg", 0.0)),
185
- fov_deg=float(a.get("fov_deg", 90.0)),
186
- ),
187
- "pan": lambda a: PanAction(delta_deg=float(a["delta_deg"])),
188
- "zoom": lambda a: ZoomAction(fov_deg=float(a["fov_deg"])),
189
- "move": lambda a: MoveAction(
190
- direction=str(a["direction"]), meters=float(a.get("meters", 10.0))
191
- ),
192
- "place_pin": lambda a: PinAction(
193
- lat=float(a["lat"]),
194
- lon=float(a["lon"]),
195
- label=a.get("label") or None,
196
- span_deg=float(a.get("span_deg", 7.0)),
197
- ),
198
- "view_map": lambda a: ViewMapAction(
199
- lat=float(a["lat"]),
200
- lon=float(a["lon"]),
201
- span_deg=float(a.get("span_deg", 7.0)),
202
- ),
203
- "measure": lambda a: MeasureAction(
204
- lat_a=float(a["lat_a"]),
205
- lon_a=float(a["lon_a"]),
206
- lat_b=float(a["lat_b"]),
207
- lon_b=float(a["lon_b"]),
208
- ),
209
- "submit_guess": lambda a: GuessAction(
210
- lat=float(a["lat"]),
211
- lon=float(a["lon"]),
212
- country=a.get("country") or None,
213
- confidence=(
214
- float(a["confidence"]) if a.get("confidence") is not None else None
215
- ),
216
- reasoning=a.get("reasoning") or None,
217
- ),
218
- }
219
-
220
-
221
- def load_tasks(
222
- index_path: str | pathlib.Path,
223
- repeat: int = 1,
224
- split: str | None = None,
225
- ) -> list[dict[str, Any]]:
226
- """
227
- Read a frozen task index into harness task dicts.
228
-
229
- Args:
230
- index_path (`str` or `pathlib.Path`):
231
- The JSONL index written by `scripts/build_tasks.py`.
232
- repeat (`int`, *optional*, defaults to `1`):
233
- Emit each task this many times consecutively. `repeat=16` gives a
234
- GRPO group of 16 rollouts per location.
235
- split (`str`, *optional*):
236
- Split name to record on every task, so the session factory selects
237
- from the right index server-side. Required whenever the server
238
- serves more than one split, because an index alone does not say
239
- which one it is.
240
-
241
- Returns:
242
- `list[dict]`: One dict per episode, each with `task_index`, `task_id`,
243
- `country` and, when given, `split`.
244
-
245
- Examples:
246
-
247
- ```python
248
- tasks = load_tasks("tasks/eval_pano_v3.jsonl", split="eval")
249
- tasks = load_tasks("tasks/train_pano_v3.jsonl", repeat=16, split="train")
250
- ```
251
- """
252
- rows = []
253
- for line in pathlib.Path(index_path).read_text().splitlines():
254
- if not line.strip():
255
- continue
256
- row = json.loads(line)
257
- task = {
258
- "task_index": int(row["task_index"]),
259
- "task_id": str(row["task_id"]),
260
- "country": str(row.get("country", "")),
261
- }
262
- if split is not None:
263
- task["split"] = split
264
- rows.extend([dict(task) for _ in range(repeat)])
265
- return rows
266
-
267
-
268
- def cycle_tasks(tasks: list[dict[str, Any]]) -> Iterator[dict[str, Any]]:
269
- """Yield tasks forever, so `num_episodes` may exceed the index size."""
270
- while True:
271
- for task in tasks:
272
- yield dict(task)
273
-
274
-
275
- def _initial_messages(result: Any, task: Any) -> list[dict[str, Any]]:
276
- """Opening message: the prompt plus the first view, as an image part."""
277
- observation = result.observation
278
- content: list[dict[str, Any]] = [{"type": "text", "text": observation.prompt}]
279
- if observation.image_base64:
280
- content.append(
281
- {
282
- "type": "image_url",
283
- "image_url": {
284
- "url": f"data:image/jpeg;base64,{observation.image_base64}"
285
- },
286
- }
287
- )
288
- return [{"role": "user", "content": content}]
289
-
290
-
291
- def _tool_result(
292
- tool_name: str, arguments: dict[str, Any], result: Any, state: Any
293
- ) -> ToolResult:
294
- """Feed the observation back as text plus, when present, an image."""
295
- observation = result.observation
296
- data: dict[str, Any] = {
297
- "feedback": observation.feedback,
298
- "heading_deg": observation.heading_deg,
299
- "fov_deg": observation.fov_deg,
300
- "steps_remaining": observation.steps_remaining,
301
- "pins": [p.model_dump() for p in observation.pins],
302
- }
303
- if observation.image_base64:
304
- data["image_base64"] = observation.image_base64
305
- data["image_kind"] = observation.image_kind
306
- if observation.distance_km is not None:
307
- data["distance_km"] = observation.distance_km
308
- return ToolResult(
309
- data=data,
310
- done=bool(result.done),
311
- metadata={
312
- "reward": result.reward,
313
- "tool": tool_name,
314
- "state": state.model_dump() if hasattr(state, "model_dump") else state,
315
- },
316
- )
317
-
318
-
319
- def _verify(
320
- transcript: list[dict[str, Any]],
321
- final_state: Any,
322
- last_result: Any,
323
- task: Any,
324
- ) -> VerifyResult:
325
- """
326
- Summarise a finished episode for the collector.
327
-
328
- `env_reward` forwards the reward the environment itself computed. Domain
329
- knowledge belongs inside the environment, so nothing here recomputes or
330
- adjusts it; the extra fields are derived statistics only.
331
- """
332
- observation = getattr(last_result, "observation", None)
333
- distance = getattr(observation, "distance_km", None)
334
- env_reward = getattr(last_result, "reward", None)
335
- metrics: dict[str, Any] = {
336
- "distance_km": float(distance) if distance is not None else -1.0,
337
- "parsed_ok": float(bool(getattr(observation, "parsed_ok", False))),
338
- "guessed": float(distance is not None),
339
- "within_200km": float(distance is not None and distance < 200.0),
340
- "within_25km": float(distance is not None and distance < 25.0),
341
- }
342
- if isinstance(task, dict) and task.get("task_index") is not None:
343
- metrics["task_index"] = float(task["task_index"])
344
- return VerifyResult(
345
- env_reward=float(env_reward) if env_reward is not None else None,
346
- done=bool(getattr(last_result, "done", False)),
347
- metrics=metrics,
348
- )
349
-
350
-
351
- class GeoGuesserSessionFactory(ResourceSessionFactory):
352
- """Create GeoGuesser-backed resource sessions for harness rollouts.
353
-
354
- Args:
355
- client_factory (`Callable[[], GeoGuesserEnv]`):
356
- Builds a client per session, so concurrent sessions stay isolated.
357
- include_navigation (`bool`, *optional*, defaults to `True`):
358
- Advertise `move` in the tool list. Set `False` for a task index
359
- whose backend cannot walk, so the model never sees a dead tool.
360
-
361
- Examples:
362
-
363
- ```python
364
- factory = GeoGuesserSessionFactory(
365
- lambda: GeoGuesserEnv(base_url="http://localhost:8000")
366
- )
367
- session = factory.create(task={"task_index": 7})
368
- ```
369
- """
370
-
371
- def __init__(
372
- self,
373
- client_factory: Callable[[], GeoGuesserEnv],
374
- *,
375
- include_navigation: bool = True,
376
- ):
377
- self._client_factory = client_factory
378
- self._tools = [
379
- tool
380
- for tool in GEOGUESSER_TOOLS
381
- if include_navigation or tool.name != "move"
382
- ]
383
-
384
- def create(
385
- self,
386
- task: Any = None,
387
- seed: int | None = None,
388
- episode_id: str | None = None,
389
- ) -> StepEnvSessionAdapter:
390
- """
391
- Open a session on one task.
392
-
393
- The task's `split` and `task_index` are passed through `reset_kwargs`,
394
- which is what makes a rollout reproducible: the same task always starts
395
- from the same panorama at the same heading. Without the split, an index
396
- is ambiguous once the server serves more than one.
397
-
398
- Args:
399
- task (`dict`, *optional*):
400
- Task dict with a `task_index` and optionally a `split`. `None`
401
- selects randomly from the server's default split.
402
- seed (`int`, *optional*):
403
- Fallback selector when no task is given.
404
- episode_id (`str`, *optional*):
405
- Episode identifier recorded by the collector.
406
-
407
- Returns:
408
- `StepEnvSessionAdapter`: The session, ready to be driven.
409
- """
410
- reset_kwargs: dict[str, Any] = {}
411
- if isinstance(task, dict):
412
- if task.get("split"):
413
- reset_kwargs["split"] = str(task["split"])
414
- if task.get("task_index") is not None:
415
- reset_kwargs["index"] = int(task["task_index"])
416
- elif isinstance(task, int):
417
- reset_kwargs["index"] = task
418
-
419
- return StepEnvSessionAdapter(
420
- client=self._client_factory(),
421
- task=task,
422
- seed=seed,
423
- episode_id=episode_id,
424
- tool_specs=list(self._tools),
425
- action_builder=lambda name, arguments: to_wire(
426
- _ACTION_BY_TOOL[name](arguments)
427
- ),
428
- initial_messages_builder=_initial_messages,
429
- tool_result_builder=_tool_result,
430
- verify_builder=_verify,
431
- reset_kwargs=reset_kwargs,
432
- )
433
-
434
-
435
- __all__ = [
436
- "GEOGUESSER_TOOLS",
437
- "GeoGuesserSessionFactory",
438
- "cycle_tasks",
439
- "load_tasks",
440
- ]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/models.py DELETED
@@ -1,462 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Data models for the GeoGuesser environment.
4
-
5
- An episode places the agent at an unknown street-level location. It may look
6
- around, walk along the road, and pin candidate coordinates on a map before
7
- committing to a final guess. Only the final guess is scored.
8
- """
9
-
10
- from __future__ import annotations
11
-
12
- from enum import Enum
13
- from typing import Any, Literal
14
-
15
- from openenv.core.env_server import Action, Observation, State
16
- from pydantic import BaseModel, Field
17
-
18
-
19
- class EpisodeMode(str, Enum):
20
- """How much of the tool surface an episode exposes.
21
-
22
- Attributes:
23
- SINGLE_SHOT:
24
- One view, one guess. No pin loop, no navigation.
25
- AGENTIC:
26
- The full tool surface, subject to what the backend supports.
27
- NMPZ:
28
- No move, pan or zoom — the competitive "NMPZ" mode. Only pinning
29
- and guessing.
30
- """
31
-
32
- SINGLE_SHOT = "single_shot"
33
- AGENTIC = "agentic"
34
- NMPZ = "nmpz"
35
-
36
-
37
- class RewardMode(str, Enum):
38
- """Which ground-truth granularity the reward is computed against."""
39
-
40
- COORDS = "coords"
41
- COUNTRY_ONLY = "country_only"
42
-
43
-
44
- # =============================================================================
45
- # Actions
46
- # =============================================================================
47
-
48
-
49
- class TypedAction(Action):
50
- """Base class for the ergonomic, per-operation action types.
51
-
52
- These are what callers construct in Python. They are converted to the flat
53
- [`GeoGuesserAction`] wire type by the client, because an HTTP environment
54
- declares exactly one action schema and `Action` forbids unknown fields.
55
- """
56
-
57
-
58
- class LookAction(TypedAction):
59
- """Render a perspective view out of the current panorama.
60
-
61
- Args:
62
- heading_deg (`float`):
63
- Absolute compass heading in degrees, `0` being true north.
64
- pitch_deg (`float`, *optional*, defaults to `0.0`):
65
- Vertical angle in degrees; positive looks up.
66
- fov_deg (`float`, *optional*, defaults to `90.0`):
67
- Horizontal field of view. Smaller values zoom in.
68
- """
69
-
70
- heading_deg: float = Field(default=0.0, ge=-3600.0, le=3600.0)
71
- pitch_deg: float = Field(default=0.0, ge=-90.0, le=90.0)
72
- fov_deg: float = Field(default=90.0, ge=10.0, le=120.0)
73
-
74
-
75
- class PanAction(TypedAction):
76
- """Turn relative to the current heading.
77
-
78
- Args:
79
- delta_deg (`float`):
80
- Degrees to turn; positive turns right.
81
- """
82
-
83
- delta_deg: float = Field(ge=-3600.0, le=3600.0)
84
-
85
-
86
- class ZoomAction(TypedAction):
87
- """Change the field of view without turning.
88
-
89
- Args:
90
- fov_deg (`float`):
91
- New horizontal field of view in degrees.
92
- """
93
-
94
- fov_deg: float = Field(ge=10.0, le=120.0)
95
-
96
-
97
- class MoveAction(TypedAction):
98
- """Walk along the captured sequence.
99
-
100
- Args:
101
- direction (`str`):
102
- Either `"forward"` or `"backward"` along the sequence.
103
- meters (`float`, *optional*, defaults to `10.0`):
104
- Requested distance. Frame spacing is irregular, so the observation
105
- reports how far the move actually travelled.
106
- """
107
-
108
- direction: str = Field(pattern="^(forward|backward)$")
109
- meters: float = Field(default=10.0, gt=0.0, le=500.0)
110
-
111
-
112
- class PinAction(TypedAction):
113
- """Place a candidate pin and receive a map of where it landed.
114
-
115
- The response describes the pinned location only. It carries no information
116
- about the true location.
117
-
118
- Args:
119
- lat (`float`):
120
- Latitude of the candidate.
121
- lon (`float`):
122
- Longitude of the candidate.
123
- label (`str`, *optional*):
124
- Free-text note carried back in the pin list.
125
- span_deg (`float`, *optional*, defaults to `7.0`):
126
- Half-width in degrees of the map window returned with the pin.
127
- Choosing the zoom matters: below roughly 4 degrees the map adds
128
- roads, urban areas and town names, which is what makes aiming
129
- within a city possible rather than guessing at its centre.
130
- """
131
-
132
- lat: float = Field(ge=-90.0, le=90.0)
133
- lon: float = Field(ge=-180.0, le=180.0)
134
- label: str | None = None
135
- span_deg: float = Field(default=7.0, gt=0.02, le=180.0)
136
-
137
-
138
- class ViewMapAction(TypedAction):
139
- """Pan and zoom the map without committing a pin.
140
-
141
- Args:
142
- lat (`float`):
143
- Latitude at the centre of the view.
144
- lon (`float`):
145
- Longitude at the centre of the view.
146
- span_deg (`float`, *optional*, defaults to `7.0`):
147
- Half-width of the window in degrees.
148
- """
149
-
150
- lat: float = Field(ge=-90.0, le=90.0)
151
- lon: float = Field(ge=-180.0, le=180.0)
152
- span_deg: float = Field(default=7.0, gt=0.05, le=180.0)
153
-
154
-
155
- class MeasureAction(TypedAction):
156
- """Great-circle distance between two of the agent's own coordinates."""
157
-
158
- lat_a: float = Field(ge=-90.0, le=90.0)
159
- lon_a: float = Field(ge=-180.0, le=180.0)
160
- lat_b: float = Field(ge=-90.0, le=90.0)
161
- lon_b: float = Field(ge=-180.0, le=180.0)
162
-
163
-
164
- class GuessAction(TypedAction):
165
- """Commit a final guess. Terminal.
166
-
167
- Either supply `response` and let the environment parse it, or supply
168
- `lat`/`lon` directly. Passing the raw reply keeps extraction failures
169
- visible in the score rather than hidden in the harness.
170
-
171
- Args:
172
- response (`str`, *optional*):
173
- The model's unedited reply. Coordinates are extracted from it.
174
- lat (`float`, *optional*):
175
- Latitude, when the caller has already parsed the reply.
176
- lon (`float`, *optional*):
177
- Longitude, when the caller has already parsed the reply.
178
- country (`str`, *optional*):
179
- ISO-3166 alpha-2 code or country name, scored for partial credit.
180
- confidence (`float`, *optional*):
181
- Self-reported confidence in [0, 1], recorded for calibration.
182
- reasoning (`str`, *optional*):
183
- Free-text rationale, recorded but not scored.
184
- """
185
-
186
- response: str | None = None
187
- lat: float | None = Field(default=None, ge=-90.0, le=90.0)
188
- lon: float | None = Field(default=None, ge=-180.0, le=180.0)
189
- country: str | None = None
190
- confidence: float | None = Field(default=None, ge=0.0, le=1.0)
191
- reasoning: str | None = None
192
-
193
-
194
- # =============================================================================
195
- # Wire action
196
- # =============================================================================
197
-
198
-
199
- class GeoGuesserAction(Action):
200
- """The single action schema the server accepts.
201
-
202
- An HTTP environment declares one action class, so every operation travels
203
- as this flat record with an `op` discriminator. Callers normally build a
204
- [`TypedAction`] subclass instead and let the client convert.
205
-
206
- Attributes:
207
- op (`str`):
208
- Which operation to perform: `"look"`, `"pan"`, `"zoom"`, `"move"`,
209
- `"pin"`, `"view_map"`, `"measure"` or `"guess"`.
210
- """
211
-
212
- op: Literal["look", "pan", "zoom", "move", "pin", "view_map", "measure", "guess"]
213
-
214
- heading_deg: float | None = None
215
- pitch_deg: float | None = None
216
- fov_deg: float | None = None
217
- delta_deg: float | None = None
218
- direction: str | None = None
219
- meters: float | None = None
220
- lat: float | None = None
221
- lon: float | None = None
222
- label: str | None = None
223
- span_deg: float | None = None
224
- lat_a: float | None = None
225
- lon_a: float | None = None
226
- lat_b: float | None = None
227
- lon_b: float | None = None
228
- response: str | None = None
229
- country: str | None = None
230
- confidence: float | None = None
231
- reasoning: str | None = None
232
-
233
-
234
- _OP_BY_TYPE: dict[type, str] = {}
235
- _TYPE_BY_OP: dict[str, type] = {}
236
-
237
-
238
- def _register_ops() -> None:
239
- pairs = [
240
- (LookAction, "look"),
241
- (PanAction, "pan"),
242
- (ZoomAction, "zoom"),
243
- (MoveAction, "move"),
244
- (PinAction, "pin"),
245
- (ViewMapAction, "view_map"),
246
- (MeasureAction, "measure"),
247
- (GuessAction, "guess"),
248
- ]
249
- for cls, op in pairs:
250
- _OP_BY_TYPE[cls] = op
251
- _TYPE_BY_OP[op] = cls
252
-
253
-
254
- _register_ops()
255
-
256
-
257
- def to_wire(action: TypedAction) -> GeoGuesserAction:
258
- """
259
- Convert a typed action into the flat wire action.
260
-
261
- Args:
262
- action ([`TypedAction`]):
263
- The action to convert.
264
-
265
- Returns:
266
- [`GeoGuesserAction`]: The same action, flattened, with `op` set.
267
- """
268
- op = _OP_BY_TYPE.get(type(action))
269
- if op is None:
270
- raise TypeError(f"No wire op registered for {type(action).__name__}")
271
- payload = action.model_dump(exclude_none=True, exclude={"metadata"})
272
- return GeoGuesserAction(op=op, **payload)
273
-
274
-
275
- def from_wire(action: GeoGuesserAction) -> TypedAction:
276
- """
277
- Rebuild the typed action a wire action stands for.
278
-
279
- Args:
280
- action ([`GeoGuesserAction`]):
281
- The received wire action.
282
-
283
- Returns:
284
- [`TypedAction`]: The corresponding typed action, validated.
285
- """
286
- cls = _TYPE_BY_OP.get(action.op)
287
- if cls is None:
288
- raise ValueError(f"Unknown op: {action.op!r}")
289
- fields = set(cls.model_fields) - {"metadata"}
290
- payload = {
291
- k: v for k, v in action.model_dump(exclude_none=True).items() if k in fields
292
- }
293
- return cls(**payload)
294
-
295
-
296
- # =============================================================================
297
- # Observation
298
- # =============================================================================
299
-
300
-
301
- class Pin(BaseModel):
302
- """One candidate pin and what the environment could say about it.
303
-
304
- Attributes:
305
- index (`int`):
306
- 1-based position in the pin list.
307
- lat (`float`):
308
- Latitude of the candidate.
309
- lon (`float`):
310
- Longitude of the candidate.
311
- label (`str` or `None`):
312
- The note the agent attached, if any.
313
- description (`str`):
314
- What the environment could say about the pinned coordinate. Never
315
- anything about the target.
316
- """
317
-
318
- index: int
319
- lat: float
320
- lon: float
321
- label: str | None = None
322
- description: str = ""
323
-
324
-
325
- class GeoGuesserObservation(Observation):
326
- """What the agent sees after a reset or a step.
327
-
328
- The schema is identical across backends. A capability the backend lacks
329
- shows up as an unregistered tool and an empty field, never as a different
330
- shape, so one policy runs against every backend.
331
-
332
- Attributes:
333
- prompt (`str`):
334
- Instructions, populated on reset.
335
- image_base64 (`str` or `None`):
336
- The most recent rendered image as base64 PNG or JPEG — a
337
- perspective view, or a map after a pin.
338
- image_kind (`str`):
339
- Either `"view"`, `"map"` or `"none"`, saying what the image shows.
340
- heading_deg (`float`):
341
- Current compass heading in degrees.
342
- pitch_deg (`float`):
343
- Current vertical angle in degrees.
344
- fov_deg (`float`):
345
- Current field of view in degrees.
346
- moved_meters (`float`):
347
- Distance actually travelled by the last move.
348
- total_moved_meters (`float`):
349
- Cumulative distance travelled this episode.
350
- can_move_forward (`bool`):
351
- Whether a forward frame exists on the sequence.
352
- can_move_backward (`bool`):
353
- Whether a backward frame exists on the sequence.
354
- available_tools (`list[str]`):
355
- Tool names this backend actually registered.
356
- steps_remaining (`int`):
357
- Actions left before the episode is cut off.
358
- pins (`list[Pin]`):
359
- Candidates placed so far, in order.
360
- feedback (`str`):
361
- Text describing the result of the last action. For a pin, this
362
- describes the pinned location and nothing about the target.
363
- captured_at (`str`):
364
- Capture date of the current panorama, `YYYY-MM` — a legitimate
365
- meta clue, as in the real game.
366
- distance_km (`float` or `None`):
367
- Distance from guess to truth. Populated only after a guess.
368
- score (`float` or `None`):
369
- Distance score in [0, 1], before action costs. After a guess only.
370
- action_cost (`float`):
371
- Reward already spent on information gathering this episode. Visible
372
- throughout, not only after the guess, so a policy can see what it
373
- has committed.
374
- true_lat (`float` or `None`):
375
- Ground truth latitude, revealed only after a guess.
376
- true_lon (`float` or `None`):
377
- Ground truth longitude, revealed only after a guess.
378
- parsed_ok (`bool`):
379
- Whether coordinates could be extracted from the guess.
380
- """
381
-
382
- prompt: str = ""
383
- image_base64: str | None = None
384
- image_kind: str = "none"
385
-
386
- heading_deg: float = 0.0
387
- pitch_deg: float = 0.0
388
- fov_deg: float = 90.0
389
-
390
- moved_meters: float = 0.0
391
- total_moved_meters: float = 0.0
392
- can_move_forward: bool = False
393
- can_move_backward: bool = False
394
-
395
- available_tools: list[str] = Field(default_factory=list)
396
- steps_remaining: int = 0
397
- pins: list[Pin] = Field(default_factory=list)
398
- feedback: str = ""
399
- captured_at: str = ""
400
-
401
- distance_km: float | None = None
402
- score: float | None = None
403
- action_cost: float | None = None
404
- true_lat: float | None = None
405
- true_lon: float | None = None
406
- parsed_ok: bool = True
407
-
408
-
409
- # =============================================================================
410
- # State
411
- # =============================================================================
412
-
413
-
414
- class GeoGuesserState(State):
415
- """Internal episode state. Never sent to the agent verbatim.
416
-
417
- Attributes:
418
- task_index (`int`):
419
- Index into the frozen task list, `-1` before the first reset.
420
- task_id (`str`):
421
- Stable identifier of the sampled task.
422
- frame_index (`int`):
423
- Position within the task's sequence.
424
- heading_deg (`float`):
425
- Current heading in degrees.
426
- pitch_deg (`float`):
427
- Current pitch in degrees.
428
- fov_deg (`float`):
429
- Current field of view in degrees.
430
- n_looks (`int`):
431
- Count of view renders, for action cost.
432
- n_maps (`int`):
433
- Count of map renders that were not pins.
434
- n_pins (`int`):
435
- Count of pins placed.
436
- n_moves (`int`):
437
- Count of moves taken.
438
- n_free (`int`):
439
- Count of free-tool calls, which cost nothing but are capped so they
440
- cannot be issued forever.
441
- total_moved_meters (`float`):
442
- Cumulative distance travelled.
443
- submitted (`bool`):
444
- Whether the single allowed guess has been made.
445
- pins (`list[dict]`):
446
- Pins placed so far.
447
- """
448
-
449
- task_index: int = -1
450
- task_id: str = ""
451
- frame_index: int = 0
452
- heading_deg: float = 0.0
453
- pitch_deg: float = 0.0
454
- fov_deg: float = 90.0
455
- n_looks: int = 0
456
- n_maps: int = 0
457
- n_pins: int = 0
458
- n_moves: int = 0
459
- n_free: int = 0
460
- total_moved_meters: float = 0.0
461
- submitted: bool = False
462
- pins: list[dict[str, Any]] = Field(default_factory=list)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/__init__.py DELETED
@@ -1,3 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Server-side implementation of the GeoGuesser environment."""
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/app.py DELETED
@@ -1,349 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """FastAPI application for the GeoGuesser environment.
4
-
5
- Configuration comes from the process environment so one image can serve every
6
- variant. `MAPILLARY_API_KEY` is needed only to fill cache misses; with a warm
7
- cache the server runs with no network access at all.
8
-
9
- Usage:
10
- uv run --project . server
11
- uvicorn server.app:app --host 0.0.0.0 --port 8000
12
- """
13
-
14
- from __future__ import annotations
15
-
16
- import inspect
17
- import logging
18
- import os
19
- import pathlib
20
-
21
- from openenv.core.env_server.http_server import create_app
22
-
23
- try:
24
- from geoguesser_env.models import GeoGuesserAction, GeoGuesserObservation
25
- from geoguesser_env.server.geoguesser_environment import GeoGuesserEnvironment
26
- except ImportError: # running uvicorn from inside envs/geoguesser_env
27
- from models import GeoGuesserAction, GeoGuesserObservation
28
-
29
- from .geoguesser_environment import GeoGuesserEnvironment
30
-
31
-
32
- logger = logging.getLogger(__name__)
33
-
34
- _ROOT = pathlib.Path(__file__).resolve().parents[1]
35
-
36
- INDEX_PATH = os.getenv("GEOGUESSER_INDEX", str(_ROOT / "tasks" / "pano_v1.jsonl"))
37
- CACHE_DIR = os.getenv("GEOGUESSER_CACHE", str(_ROOT / "data" / "panos"))
38
- # Named splits, as (environment variable, repo-relative default). Each is
39
- # optional and a split whose file is absent is simply not offered, so the same
40
- # image serves a checkout with the indexes committed, a Space with a bucket
41
- # mounted at /data, and a deployment carrying only the eval set.
42
- SPLIT_SOURCES = {
43
- "train": ("GEOGUESSER_TASKS_TRAIN", "tasks/train_pano_v3.jsonl"),
44
- "eval": ("GEOGUESSER_TASKS_EVAL", "tasks/eval_pano_v3.jsonl"),
45
- }
46
- DEFAULT_SPLIT = os.getenv("GEOGUESSER_DEFAULT_SPLIT", "")
47
- EPISODE_MODE = os.getenv("GEOGUESSER_EPISODE_MODE", "agentic")
48
- REWARD_MODE = os.getenv("GEOGUESSER_REWARD_MODE", "coords")
49
- MAX_STEPS = int(os.getenv("GEOGUESSER_MAX_STEPS", "24"))
50
- VIEW_SIZE = int(os.getenv("GEOGUESSER_VIEW_SIZE", "640"))
51
- HIERARCHICAL = os.getenv("GEOGUESSER_HIERARCHICAL", "0") in {"1", "true", "True"}
52
- # Training defaults differ from play defaults on purpose; see the reward-shape
53
- # note in the README. A hosted Space serves play traffic, so it keeps the game
54
- # curve and shows the licence credit unless these are set explicitly.
55
- REWARD_SHAPE = os.getenv("GEOGUESSER_REWARD_SHAPE", "geoguessr")
56
- COST_MODE = os.getenv("GEOGUESSER_COST_MODE", "subtract")
57
- HIDE_IDENTITY = os.getenv("GEOGUESSER_HIDE_IDENTITY", "0") in {"1", "true", "True"}
58
- ALLOW_FETCH = os.getenv("GEOGUESSER_ALLOW_FETCH", "1") in {"1", "true", "True"}
59
- HIRES_ZOOM = os.getenv("GEOGUESSER_HIRES_ZOOM", "1") in {"1", "true", "True"}
60
- REVEAL_MAP = os.getenv("GEOGUESSER_REVEAL_MAP", "1") in {"1", "true", "True"}
61
- STREET_DETAIL = os.getenv("GEOGUESSER_STREET_DETAIL", "1") in {"1", "true", "True"}
62
- MAX_CONCURRENT = int(os.getenv("MAX_CONCURRENT_ENVS", "4"))
63
-
64
-
65
- try:
66
- from .render.minimap import set_street_detail
67
- except ImportError: # running uvicorn from inside envs/geoguesser_env
68
- from render.minimap import set_street_detail
69
-
70
- # Deliberately independent of ALLOW_FETCH. That flag governs Mapillary imagery,
71
- # which a mirrored dataset must never reach for; street detail comes from
72
- # Overpass and caches to the container's own writable disk, so a fully mirrored
73
- # deployment can still draw labelled streets. Coupling the two silently gave a
74
- # Space unlabelled agent maps while a local run had labelled ones.
75
- set_street_detail(STREET_DETAIL)
76
-
77
-
78
- def resolve_splits() -> tuple[dict[str, str], str]:
79
- """
80
- Work out which named splits this deployment actually serves.
81
-
82
- A split is offered only when its environment variable is set *and* the file
83
- exists, because a Space that mounts a bucket read-only should degrade to the
84
- splits it really has rather than failing to start. When none are configured
85
- the legacy single `GEOGUESSER_INDEX` becomes one `train` split, so existing
86
- containers behave exactly as before.
87
-
88
- Returns:
89
- `tuple` of:
90
- - `dict[str, str]`: Split name to index path.
91
- - `str`: Name of the default split.
92
- """
93
- splits: dict[str, str] = {}
94
- for name, (variable, relative) in SPLIT_SOURCES.items():
95
- override = os.getenv(variable)
96
- path = override or str(_ROOT / relative)
97
- if not pathlib.Path(path).exists():
98
- # Only complain when someone asked for it explicitly. A missing
99
- # default just means this checkout has not built that split yet.
100
- if override:
101
- logger.warning(
102
- "%s points at %s, which does not exist; "
103
- "the %r split will not be offered",
104
- variable,
105
- path,
106
- name,
107
- )
108
- continue
109
- splits[name] = path
110
-
111
- if not splits:
112
- return {"train": INDEX_PATH}, "train"
113
-
114
- default = DEFAULT_SPLIT or ("train" if "train" in splits else next(iter(splits)))
115
- if default not in splits:
116
- logger.warning(
117
- "GEOGUESSER_DEFAULT_SPLIT=%r is not among %s; using %r",
118
- DEFAULT_SPLIT,
119
- sorted(splits),
120
- next(iter(splits)),
121
- )
122
- default = next(iter(splits))
123
- return splits, default
124
-
125
-
126
- SPLITS, ACTIVE_DEFAULT_SPLIT = resolve_splits()
127
-
128
-
129
- def create_geoguesser_environment() -> GeoGuesserEnvironment:
130
- """Factory: a fresh environment per WebSocket session."""
131
- return GeoGuesserEnvironment(
132
- splits=SPLITS,
133
- default_split=ACTIVE_DEFAULT_SPLIT,
134
- cache_dir=CACHE_DIR,
135
- episode_mode=EPISODE_MODE,
136
- max_steps=MAX_STEPS,
137
- reward_mode=REWARD_MODE,
138
- hierarchical_reward=HIERARCHICAL,
139
- reward_shape=REWARD_SHAPE,
140
- cost_mode=COST_MODE,
141
- hide_task_identity=HIDE_IDENTITY,
142
- view_size=VIEW_SIZE,
143
- allow_fetch=ALLOW_FETCH,
144
- hires_zoom=HIRES_ZOOM,
145
- reveal_map=REVEAL_MAP,
146
- )
147
-
148
-
149
- def _build_app():
150
- """Create the app, attaching the Gradio tab when this openenv supports it."""
151
- kwargs = dict(
152
- env_name="geoguesser_env",
153
- max_concurrent_envs=MAX_CONCURRENT,
154
- )
155
- signature = inspect.signature(create_app)
156
- # Land people on the game, not the raw action form.
157
- if "custom_tab_primary" in signature.parameters:
158
- kwargs["custom_tab_primary"] = True
159
- if "custom_tab_name" in signature.parameters:
160
- kwargs["custom_tab_name"] = "Try Environment"
161
- if "default_tab_name" in signature.parameters:
162
- kwargs["default_tab_name"] = "MCP Playground"
163
- if "title_override" in signature.parameters:
164
- kwargs["title_override"] = "Geoguesser Environment"
165
- if "gradio_builder" in signature.parameters:
166
- try:
167
- from .gradio_ui import build_geoguesser_gradio_app
168
-
169
- kwargs["gradio_builder"] = build_geoguesser_gradio_app
170
- except Exception as exc: # pragma: no cover - optional UI dependency
171
- logger.warning("Gradio UI unavailable: %r", exc)
172
- else:
173
- logger.warning(
174
- "Installed openenv does not support gradio_builder; "
175
- "the GeoGuessr-style play tab will not be available."
176
- )
177
- return create_app(
178
- create_geoguesser_environment,
179
- GeoGuesserAction,
180
- GeoGuesserObservation,
181
- **kwargs,
182
- )
183
-
184
-
185
- def _attach_play_routes(application) -> None:
186
- """Serve panoramas and task metadata to the browser-side viewers.
187
-
188
- The Pannellum viewer needs the raw equirectangular JPEG, which the agent
189
- never receives — it only ever sees reprojected views. Ground truth is
190
- exposed here because these routes exist for a human playing a round in
191
- their own browser; the agent's observations still withhold it until it
192
- guesses.
193
- """
194
- from fastapi import HTTPException, Query
195
- from fastapi.responses import FileResponse, HTMLResponse, JSONResponse
196
-
197
- # One environment, reused for metadata only. Its per-split backends share
198
- # the process-wide parsed index cache, so this is cheap.
199
- catalog = create_geoguesser_environment()
200
-
201
- def _backend(split: str | None):
202
- """Resolve a split name to its backend, as a 404 rather than a 500."""
203
- try:
204
- return catalog._backend_for(split or ACTIVE_DEFAULT_SPLIT)
205
- except KeyError as exc:
206
- raise HTTPException(status_code=404, detail=str(exc)) from exc
207
-
208
- @application.get(
209
- "/geoguesser/task/{task_index}",
210
- tags=["geoguesser"],
211
- response_class=JSONResponse,
212
- )
213
- async def geoguesser_task(task_index: int, split: str | None = Query(None)):
214
- """Metadata for one task, for the play UI."""
215
- try:
216
- task = _backend(split).task(task_index)
217
- except IndexError as exc:
218
- raise HTTPException(status_code=404, detail=str(exc)) from exc
219
- frame = task.frames[task.start_frame]
220
- return JSONResponse(
221
- {
222
- "task_index": task.task_index,
223
- "task_id": task.task_id,
224
- "lat": frame.lat,
225
- "lon": frame.lon,
226
- "country": task.country,
227
- "compass_angle": frame.compass_angle,
228
- "captured_at": frame.captured_at,
229
- "attribution": task.attribution,
230
- "n_frames": len(task.frames),
231
- "start_frame": task.start_frame,
232
- # Per-frame headings only. Coordinates are withheld for frames
233
- # other than the start, which is the one the guess is scored
234
- # against and therefore already revealed to a human player.
235
- "frames": [
236
- {
237
- "index": i,
238
- "compass_angle": f.compass_angle,
239
- "captured_at": f.captured_at,
240
- }
241
- for i, f in enumerate(task.frames)
242
- ],
243
- }
244
- )
245
-
246
- def _pano_response(task_index: int, frame_index: int | None, split: str | None):
247
- """Resolve one frame's panorama file, fetching it if necessary."""
248
- resolved = _backend(split)
249
- try:
250
- task = resolved.task(task_index)
251
- except IndexError as exc:
252
- raise HTTPException(status_code=404, detail=str(exc)) from exc
253
- position = task.start_frame if frame_index is None else frame_index
254
- if not 0 <= position < len(task.frames):
255
- raise HTTPException(
256
- status_code=404,
257
- detail=(
258
- f"frame {position} out of range for task {task_index} "
259
- f"with {len(task.frames)} frames"
260
- ),
261
- )
262
- frame = task.frames[position]
263
- try:
264
- resolved.load_pano(frame.image_id)
265
- except Exception as exc:
266
- raise HTTPException(status_code=503, detail=str(exc)) from exc
267
- return FileResponse(
268
- pathlib.Path(CACHE_DIR) / f"{frame.image_id}.jpg",
269
- media_type="image/jpeg",
270
- )
271
-
272
- @application.get(
273
- "/geoguesser/pano/{task_index}",
274
- tags=["geoguesser"],
275
- response_class=FileResponse,
276
- )
277
- async def geoguesser_pano(task_index: int, split: str | None = Query(None)):
278
- """The task's starting panorama, for the browser viewer."""
279
- return _pano_response(task_index, None, split)
280
-
281
- @application.get(
282
- "/geoguesser/pano/{task_index}/{frame_index}",
283
- tags=["geoguesser"],
284
- response_class=FileResponse,
285
- )
286
- async def geoguesser_pano_frame(
287
- task_index: int, frame_index: int, split: str | None = Query(None)
288
- ):
289
- """One specific frame's panorama, so the viewer can follow `move()`."""
290
- return _pano_response(task_index, frame_index, split)
291
-
292
- @application.get(
293
- "/geoguesser/play", tags=["geoguesser"], response_class=HTMLResponse
294
- )
295
- async def geoguesser_play(split: str | None = Query(None)):
296
- """The standalone play page, also embedded in the Gradio tab."""
297
- from .gradio_ui import play_page_html
298
-
299
- return HTMLResponse(play_page_html(catalog.list_splits(), split))
300
-
301
- @application.get(
302
- "/geoguesser/tasks", tags=["geoguesser"], response_class=JSONResponse
303
- )
304
- async def geoguesser_tasks():
305
- """Which splits exist and how many tasks each holds."""
306
- splits = catalog.list_splits()
307
- return JSONResponse(
308
- {
309
- "splits": splits,
310
- "default_split": ACTIVE_DEFAULT_SPLIT,
311
- # Kept so an older play page still finds a count.
312
- "n_tasks": next(
313
- s["num_tasks"] for s in splits if s["name"] == ACTIVE_DEFAULT_SPLIT
314
- ),
315
- }
316
- )
317
-
318
-
319
- app = _build_app()
320
-
321
- try:
322
- _attach_play_routes(app)
323
- except Exception as exc: # pragma: no cover - index may be absent in CI
324
- logger.warning("play routes unavailable: %r", exc)
325
-
326
-
327
- def main(host: str = "0.0.0.0", port: int = 8000) -> None:
328
- """
329
- Entry point for running the server without Docker.
330
-
331
- Args:
332
- host (`str`, *optional*, defaults to `"0.0.0.0"`):
333
- Address to bind.
334
- port (`int`, *optional*, defaults to `8000`):
335
- Port to listen on.
336
- """
337
- import uvicorn
338
-
339
- uvicorn.run(app, host=host, port=port)
340
-
341
-
342
- if __name__ == "__main__":
343
- import argparse
344
-
345
- parser = argparse.ArgumentParser()
346
- parser.add_argument("--port", type=int, default=8000)
347
- parser.add_argument("--host", default="0.0.0.0")
348
- args = parser.parse_args()
349
- main(host=args.host, port=args.port)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/backends/__init__.py DELETED
@@ -1,7 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Imagery backends for the GeoGuesser environment."""
4
-
5
- from .base import Frame, PanoramaBackend, Task
6
-
7
- __all__ = ["Frame", "PanoramaBackend", "Task"]
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/backends/base.py DELETED
@@ -1,127 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """The contract every imagery backend implements.
4
-
5
- Backends differ only in where pixels come from and whether the location can be
6
- walked. Everything above them — scoring, parsing, the pin loop, the map — is
7
- shared, so a policy trained against one backend runs unmodified against
8
- another.
9
-
10
- A backend that cannot do something reports it through `supports_look` or
11
- `supports_move`; the environment then simply does not register the
12
- corresponding tools. Registering a tool that always fails would only teach a
13
- policy to spend its step budget discovering that.
14
- """
15
-
16
- from __future__ import annotations
17
-
18
- from dataclasses import dataclass, field
19
- from typing import Any, Protocol, runtime_checkable
20
-
21
- from PIL import Image
22
-
23
-
24
- @dataclass(frozen=True)
25
- class Frame:
26
- """One captured position within a sequence.
27
-
28
- Attributes:
29
- image_id (`str`):
30
- Provider-side identifier, also the cache filename.
31
- lat (`float`):
32
- Latitude in degrees.
33
- lon (`float`):
34
- Longitude in degrees.
35
- compass_angle (`float`):
36
- Heading the camera faced, in degrees clockwise from true north.
37
- captured_at (`str`):
38
- Capture month as `YYYY-MM`.
39
- is_pano (`bool`, *optional*, defaults to `True`):
40
- Whether this frame is a full 360-degree panorama.
41
- """
42
-
43
- image_id: str
44
- lat: float
45
- lon: float
46
- compass_angle: float
47
- captured_at: str
48
- is_pano: bool = True
49
-
50
-
51
- @dataclass(frozen=True)
52
- class Task:
53
- """One episode's location, frozen at index build time.
54
-
55
- Attributes:
56
- task_index (`int`):
57
- Position in the task list. Stable, and what `reset` selects on.
58
- task_id (`str`):
59
- Human-readable identifier, unique within the index.
60
- frames (`list[Frame]`):
61
- Ordered frames of the captured sequence, walkable with `move`.
62
- start_frame (`int`):
63
- Index into `frames` where the episode begins.
64
- country (`str`):
65
- ISO-3166 alpha-2 code of the true location. Ground truth — never
66
- placed in an observation before the guess.
67
- sequence_id (`str`):
68
- Provider-side sequence identifier.
69
- provider (`str`, *optional*, defaults to `"mapillary"`):
70
- Which backend can resolve this task's imagery.
71
- attribution (`dict`, *optional*):
72
- Creator credit, required by the CC-BY-SA licence on the imagery.
73
- meta (`dict`, *optional*):
74
- Anything else the builder recorded — camera make and model,
75
- quality score, checksums.
76
- """
77
-
78
- task_index: int
79
- task_id: str
80
- frames: list[Frame]
81
- start_frame: int
82
- country: str
83
- sequence_id: str
84
- provider: str = "mapillary"
85
- attribution: dict[str, Any] = field(default_factory=dict)
86
- meta: dict[str, Any] = field(default_factory=dict)
87
-
88
- @property
89
- def truth(self) -> tuple[float, float]:
90
- """Ground-truth `(lat, lon)` of the starting frame."""
91
- frame = self.frames[self.start_frame]
92
- return frame.lat, frame.lon
93
-
94
-
95
- @runtime_checkable
96
- class PanoramaBackend(Protocol):
97
- """Resolves tasks to imagery and answers what the agent may do."""
98
-
99
- @property
100
- def n_tasks(self) -> int:
101
- """Number of tasks in the frozen index."""
102
- ...
103
-
104
- @property
105
- def supports_look(self) -> bool:
106
- """Whether views can be rendered at arbitrary headings."""
107
- ...
108
-
109
- @property
110
- def supports_move(self) -> bool:
111
- """Whether the location can be walked along a sequence."""
112
- ...
113
-
114
- def task(self, task_index: int) -> Task:
115
- """Return the task at `task_index`."""
116
- ...
117
-
118
- def render_view(
119
- self,
120
- task: Task,
121
- frame_index: int,
122
- heading_deg: float,
123
- pitch_deg: float,
124
- fov_deg: float,
125
- ) -> Image.Image:
126
- """Render what the camera sees from one frame."""
127
- ...
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/backends/panorama.py DELETED
@@ -1,399 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Mapillary-backed panoramas, read from a frozen index and a local cache.
4
-
5
- The index (`tasks/pano_v1.jsonl`) is self-contained: it carries every frame's
6
- coordinates, heading and capture date, so the movement graph resolves with no
7
- network access at all. Only image *bytes* may need fetching, and only on a
8
- cache miss, because Mapillary's `thumb_*_url` values are expiring signed CDN
9
- URLs and cannot be stored in the index.
10
-
11
- Once a task's frames are cached, episodes are byte-identical on repeat — the
12
- property a GRPO group depends on.
13
- """
14
-
15
- from __future__ import annotations
16
-
17
- import hashlib
18
- import json
19
- import logging
20
- import os
21
- import pathlib
22
- import urllib.parse
23
- import urllib.request
24
-
25
- from PIL import Image
26
-
27
- from ..render.pano import look
28
- from .base import Frame, Task
29
-
30
-
31
- logger = logging.getLogger(__name__)
32
-
33
- GRAPH_API = "https://graph.mapillary.com"
34
- FETCH_TIMEOUT_S = 60.0
35
-
36
- # Zooming into a 2048x1024 panorama is resolution-starved: a 30-degree view
37
- # samples only ~170 source pixels, so narrowing the field of view barely adds
38
- # detail (measured mean gradient 6.60 at 90 degrees against 7.03 at 30). The
39
- # 7680x3840 original roughly doubles it instead (10.23 against 14.87), which is
40
- # what makes reading a distant sign possible at all.
41
- #
42
- # So two derivatives are cached per panorama: the 2048 for wide views, which
43
- # renders in ~30 ms, and the original for zoomed views at ~70 ms and 3.3 MB.
44
- THUMB_VARIANT = "thumb_2048_url"
45
- ORIGINAL_VARIANT = "thumb_original_url"
46
- HIRES_FOV_DEG = 45.0
47
-
48
-
49
- class MissingImageError(RuntimeError):
50
- """A frame is absent from the cache and cannot be fetched."""
51
-
52
-
53
- _INDEX_CACHE: dict[tuple[str, float, int], list[Task]] = {}
54
- """Parsed indexes keyed by path, mtime and size, so edits invalidate the entry."""
55
-
56
-
57
- def load_index(index_path: str | pathlib.Path) -> list[Task]:
58
- """
59
- Parse a task index, reusing an already-parsed copy when possible.
60
-
61
- The Task API routes construct a throwaway environment per request and a
62
- 5,000-task index is roughly 31 MB, so re-parsing it on every call is not
63
- affordable. Keying on mtime and size means rebuilding an index in place
64
- invalidates the entry instead of serving stale tasks.
65
-
66
- Args:
67
- index_path (`str` or `pathlib.Path`):
68
- JSONL index written by `scripts/build_tasks.py`.
69
-
70
- Returns:
71
- `list[Task]`: Tasks ordered by `task_index`. The list is shared between
72
- callers, so treat it as read-only.
73
-
74
- Raises:
75
- FileNotFoundError: When the index does not exist.
76
- ValueError: When the index is empty, or its task indices are not
77
- contiguous from zero.
78
- """
79
- path = pathlib.Path(index_path)
80
- if not path.exists():
81
- raise FileNotFoundError(
82
- f"Task index not found: {path}. Build one with scripts/build_tasks.py."
83
- )
84
- stat = path.stat()
85
- key = (str(path.resolve()), stat.st_mtime, stat.st_size)
86
- cached = _INDEX_CACHE.get(key)
87
- if cached is not None:
88
- return cached
89
-
90
- tasks: list[Task] = []
91
- for line in path.read_text().splitlines():
92
- line = line.strip()
93
- if not line:
94
- continue
95
- row = json.loads(line)
96
- frames = [
97
- Frame(
98
- image_id=str(f["image_id"]),
99
- lat=float(f["lat"]),
100
- lon=float(f["lon"]),
101
- compass_angle=float(f.get("compass_angle", 0.0)),
102
- captured_at=str(f.get("captured_at", "")),
103
- is_pano=bool(f.get("is_pano", True)),
104
- )
105
- for f in row["frames"]
106
- ]
107
- tasks.append(
108
- Task(
109
- task_index=int(row["task_index"]),
110
- task_id=str(row["task_id"]),
111
- frames=frames,
112
- start_frame=int(row.get("start_frame", 0)),
113
- country=str(row.get("country", "")),
114
- sequence_id=str(row.get("sequence_id", "")),
115
- provider=str(row.get("provider", "mapillary")),
116
- attribution=row.get("attribution", {}),
117
- meta=row.get("meta", {}),
118
- )
119
- )
120
- if not tasks:
121
- raise ValueError(f"Task index {path} is empty.")
122
- tasks.sort(key=lambda t: t.task_index)
123
- for position, task in enumerate(tasks):
124
- if task.task_index != position:
125
- raise ValueError(
126
- "Task indices must be contiguous from 0; found "
127
- f"{task.task_index} at position {position}."
128
- )
129
- _INDEX_CACHE[key] = tasks
130
- return tasks
131
-
132
-
133
- class PanoramaBackend:
134
- """Serve panoramas from a frozen task index plus a disk cache.
135
-
136
- Args:
137
- index_path (`str` or `pathlib.Path`):
138
- JSONL task index produced by `scripts/build_pano_tasks.py`.
139
- cache_dir (`str` or `pathlib.Path`):
140
- Directory holding cached JPEGs, named `<image_id>.jpg`.
141
- access_token (`str`, *optional*):
142
- Mapillary token, used only to fill cache misses. When absent, a
143
- miss raises [`MissingImageError`] instead of reaching the network.
144
- allow_fetch (`bool`, *optional*, defaults to `True`):
145
- Set `False` to guarantee an episode never touches the network.
146
- hires_zoom (`bool`, *optional*, defaults to `True`):
147
- Render fields of view at or below `hires_fov_deg` from the
148
- full-resolution original, so zooming actually resolves detail.
149
- Falls back to the 2048 derivative when no original exists.
150
- hires_fov_deg (`float`, *optional*, defaults to `45.0`):
151
- Field of view at or below which the original is used.
152
- verify_checksums (`bool`, *optional*, defaults to `False`):
153
- Verify each cached start frame against the sha256 recorded at build
154
- time. Used by frozen evals to detect drift.
155
-
156
- Examples:
157
-
158
- ```python
159
- backend = PanoramaBackend("tasks/pano_v1.jsonl", "data/panos")
160
- task = backend.task(0)
161
- view = backend.render_view(task, task.start_frame, 90.0, 0.0, 90.0)
162
- ```
163
- """
164
-
165
- def __init__(
166
- self,
167
- index_path: str | pathlib.Path,
168
- cache_dir: str | pathlib.Path,
169
- access_token: str | None = None,
170
- allow_fetch: bool = True,
171
- verify_checksums: bool = False,
172
- hires_zoom: bool = True,
173
- hires_fov_deg: float = HIRES_FOV_DEG,
174
- ):
175
- self._index_path = pathlib.Path(index_path)
176
- self._cache_dir = pathlib.Path(cache_dir)
177
- self._cache_dir.mkdir(parents=True, exist_ok=True)
178
- self._token = access_token or os.environ.get("MAPILLARY_API_KEY")
179
- self._allow_fetch = allow_fetch
180
- self._verify_checksums = verify_checksums
181
- self._hires_zoom = hires_zoom
182
- self._hires_fov_deg = hires_fov_deg
183
- self._tasks = self._load_index()
184
- logger.info(
185
- "loaded %d tasks from %s (cache: %s, fetch: %s)",
186
- len(self._tasks),
187
- self._index_path,
188
- self._cache_dir,
189
- "on" if self._allow_fetch and self._token else "off",
190
- )
191
-
192
- # -- index ------------------------------------------------------------
193
-
194
- def _load_index(self) -> list[Task]:
195
- return load_index(self._index_path)
196
-
197
- # -- capabilities -----------------------------------------------------
198
-
199
- @property
200
- def n_tasks(self) -> int:
201
- """Number of tasks in the frozen index."""
202
- return len(self._tasks)
203
-
204
- @property
205
- def supports_look(self) -> bool:
206
- """True — every task in this index is a 360-degree panorama."""
207
- return True
208
-
209
- @property
210
- def supports_move(self) -> bool:
211
- """Whether any task has more than one frame to walk between."""
212
- return any(len(t.frames) > 1 for t in self._tasks)
213
-
214
- def task(self, task_index: int) -> Task:
215
- """
216
- Return the task at `task_index`.
217
-
218
- Args:
219
- task_index (`int`):
220
- Position in the frozen index.
221
-
222
- Returns:
223
- [`Task`]: The task, including its full frame list.
224
- """
225
- if not 0 <= task_index < len(self._tasks):
226
- raise IndexError(
227
- f"task_index {task_index} out of range for {len(self._tasks)} tasks."
228
- )
229
- return self._tasks[task_index]
230
-
231
- # -- imagery ----------------------------------------------------------
232
-
233
- def _cache_path(self, image_id: str, hires: bool = False) -> pathlib.Path:
234
- suffix = ".orig.jpg" if hires else ".jpg"
235
- return self._cache_dir / f"{image_id}{suffix}"
236
-
237
- def _fetch(self, image_id: str, hires: bool = False) -> bytes:
238
- if not self._allow_fetch:
239
- raise MissingImageError(
240
- f"Image {image_id} is not cached and fetching is disabled. "
241
- "Warm the cache with scripts/build_pano_tasks.py --prefetch-frames."
242
- )
243
- if not self._token:
244
- raise MissingImageError(
245
- f"Image {image_id} is not cached and MAPILLARY_API_KEY is unset, "
246
- "so it cannot be fetched."
247
- )
248
- variant = ORIGINAL_VARIANT if hires else THUMB_VARIANT
249
- meta_url = f"{GRAPH_API}/{image_id}?" + urllib.parse.urlencode(
250
- {"fields": variant, "access_token": self._token}
251
- )
252
- with urllib.request.urlopen(meta_url, timeout=FETCH_TIMEOUT_S) as response:
253
- thumb_url = json.loads(response.read()).get(variant)
254
- if not thumb_url:
255
- raise MissingImageError(
256
- f"Mapillary returned no {variant} for {image_id}; its "
257
- "derivatives may have been removed."
258
- )
259
- with urllib.request.urlopen(thumb_url, timeout=FETCH_TIMEOUT_S) as response:
260
- return response.read()
261
-
262
- def load_pano(
263
- self,
264
- image_id: str,
265
- expected_sha256: str | None = None,
266
- hires: bool = False,
267
- ) -> Image.Image:
268
- """
269
- Return a panorama, fetching and caching it if necessary.
270
-
271
- Args:
272
- image_id (`str`):
273
- Provider-side image identifier.
274
- expected_sha256 (`str`, *optional*):
275
- Checksum recorded at build time. Verified only when the backend
276
- was constructed with `verify_checksums=True`, and only for
277
- the 2048 derivative, which is what the index records.
278
- hires (`bool`, *optional*, defaults to `False`):
279
- Load the full-resolution original instead of the 2048
280
- derivative.
281
-
282
- Returns:
283
- `PIL.Image.Image`: The equirectangular panorama.
284
- """
285
- path = self._cache_path(image_id, hires=hires)
286
- if not path.exists():
287
- payload = self._fetch(image_id, hires=hires)
288
- path.write_bytes(payload)
289
- logger.info(
290
- "cached %s%s (%.0f KB)",
291
- image_id,
292
- " at full resolution" if hires else "",
293
- len(payload) / 1024,
294
- )
295
- if self._verify_checksums and expected_sha256 and not hires:
296
- actual = hashlib.sha256(path.read_bytes()).hexdigest()
297
- if actual != expected_sha256:
298
- raise MissingImageError(
299
- f"Checksum mismatch for {image_id}: index recorded "
300
- f"{expected_sha256[:12]}, cache holds {actual[:12]}. The "
301
- "upstream image changed; this task is no longer comparable."
302
- )
303
- return Image.open(path)
304
-
305
- def render_view(
306
- self,
307
- task: Task,
308
- frame_index: int,
309
- heading_deg: float,
310
- pitch_deg: float = 0.0,
311
- fov_deg: float = 90.0,
312
- ) -> Image.Image:
313
- """
314
- Render what the camera sees from one frame of a task.
315
-
316
- Headings are absolute: `0` is true north, obtained by offsetting the
317
- request by the frame's own `compass_angle`. That keeps `look(0)`
318
- meaning the same thing in every task.
319
-
320
- Args:
321
- task ([`Task`]):
322
- The task being played.
323
- frame_index (`int`):
324
- Which frame of the sequence the agent stands on.
325
- heading_deg (`float`):
326
- Absolute compass heading in degrees.
327
- pitch_deg (`float`, *optional*, defaults to `0.0`):
328
- Vertical angle in degrees.
329
- fov_deg (`float`, *optional*, defaults to `90.0`):
330
- Horizontal field of view in degrees.
331
-
332
- Returns:
333
- `PIL.Image.Image`: The rendered view.
334
- """
335
- frame = task.frames[frame_index]
336
- checksums = task.meta.get("sha256", {})
337
- want_hires = self._hires_zoom and fov_deg <= self._hires_fov_deg
338
- try:
339
- pano = self.load_pano(
340
- frame.image_id, checksums.get(frame.image_id), hires=want_hires
341
- )
342
- except MissingImageError:
343
- if not want_hires:
344
- raise
345
- # A missing original must not end an episode; a soft view beats a
346
- # failed step.
347
- logger.warning(
348
- "no full-resolution original for %s; zooming on the 2048 "
349
- "derivative instead",
350
- frame.image_id,
351
- )
352
- pano = self.load_pano(frame.image_id, checksums.get(frame.image_id))
353
- return look(pano, heading_deg + frame.compass_angle, pitch_deg, fov_deg)
354
-
355
- # -- navigation -------------------------------------------------------
356
-
357
- def step_along(
358
- self, task: Task, frame_index: int, direction: str, meters: float
359
- ) -> tuple[int, float]:
360
- """
361
- Walk the sequence and report where the agent actually ended up.
362
-
363
- Frame spacing is irregular — measured around 3.3 m on Mapillary
364
- sequences — so the requested distance is consumed frame by frame and
365
- the realised distance is returned rather than assumed.
366
-
367
- Args:
368
- task ([`Task`]):
369
- The task being played.
370
- frame_index (`int`):
371
- Current position in `task.frames`.
372
- direction (`str`):
373
- `"forward"` or `"backward"`.
374
- meters (`float`):
375
- Requested distance in metres.
376
-
377
- Returns:
378
- `tuple[int, float]` with:
379
- - the new frame index, unchanged at a dead end
380
- - metres actually travelled
381
- """
382
- from ..scoring import haversine_km
383
-
384
- step = 1 if direction == "forward" else -1
385
- current = frame_index
386
- travelled = 0.0
387
- while travelled < meters:
388
- nxt = current + step
389
- if not 0 <= nxt < len(task.frames):
390
- break
391
- a, b = task.frames[current], task.frames[nxt]
392
- travelled += haversine_km(a.lat, a.lon, b.lat, b.lon) * 1000.0
393
- current = nxt
394
- return current, travelled
395
-
396
- def can_move(self, task: Task, frame_index: int, direction: str) -> bool:
397
- """Whether a frame exists in `direction` from the current position."""
398
- step = 1 if direction == "forward" else -1
399
- return 0 <= frame_index + step < len(task.frames)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/geoguesser_environment.py DELETED
@@ -1,998 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """The GeoGuesser environment.
4
-
5
- Exposes its tools over MCP so an agent can look around, walk, and check
6
- candidate coordinates on a map before committing to a guess. Non-MCP
7
- structured actions route through `_step_impl`, which is what makes single-shot
8
- GRPO and the agentic loop the same environment rather than two.
9
-
10
- Two rules are load-bearing:
11
-
12
- - Pin feedback describes only where the agent pointed. Any signal about the
13
- target would make binary search optimal, and the benchmark would measure
14
- bisection instead of geography.
15
- - Tools the backend cannot serve are never registered, rather than registered
16
- and failing, so a policy does not learn to spend steps on dead ends.
17
- """
18
-
19
- from __future__ import annotations
20
-
21
- import logging
22
- import random
23
- import uuid
24
- from typing import Any
25
-
26
- from fastmcp import FastMCP
27
- from openenv.core.env_server.mcp_environment import MCPEnvironment
28
- from openenv.core.env_server.types import Action, Observation
29
-
30
- from ..models import (
31
- EpisodeMode,
32
- from_wire,
33
- GeoGuesserAction,
34
- GeoGuesserObservation,
35
- GeoGuesserState,
36
- GuessAction,
37
- LookAction,
38
- MeasureAction,
39
- MoveAction,
40
- PanAction,
41
- Pin,
42
- PinAction,
43
- RewardMode,
44
- TypedAction,
45
- ViewMapAction,
46
- ZoomAction,
47
- )
48
- from .backends.base import Task
49
- from .backends.panorama import PanoramaBackend
50
- from .parser import parse_guess
51
- from .render.minimap import (
52
- describe_pin,
53
- locate,
54
- render_map,
55
- street_detail_enabled,
56
- street_fetch_failed,
57
- )
58
- from .render.pano import to_base64
59
- from .scoring import action_cost, compute_reward, haversine_km, verdict
60
-
61
-
62
- logger = logging.getLogger(__name__)
63
-
64
- PROMPT = (
65
- "You are dropped at an unknown street-level location somewhere in the "
66
- "world. Work out where you are.\n\n"
67
- "Available tools: {tools}\n\n"
68
- "Looking around and checking the map cost a little reward each, so gather "
69
- "the evidence you need and then commit. You have {steps} actions. "
70
- "Placing a pin shows you where on the map that coordinate falls - it "
71
- "tells you nothing about whether you are right. Finish with "
72
- "submit_guess."
73
- )
74
-
75
-
76
- class UnknownSplitError(KeyError, IndexError):
77
- """A split name that is not configured.
78
-
79
- Inherits both exception types deliberately. `KeyError` is what a Python
80
- caller expects from a bad name, while the core Task API dispatcher maps
81
- only `NotImplementedError` and `IndexError` onto HTTP status codes -- so
82
- without `IndexError` an unknown split surfaces as a 500 instead of a 400.
83
- """
84
-
85
-
86
- class GeoGuesserEnvironment(MCPEnvironment):
87
- """A GeoGuessr-style geolocation episode.
88
-
89
- Args:
90
- index_path (`str`):
91
- JSONL task index built by `scripts/build_pano_tasks.py`.
92
- cache_dir (`str`):
93
- Directory of cached panorama JPEGs.
94
- episode_mode (`str`, *optional*, defaults to `"agentic"`):
95
- One of `"agentic"`, `"single_shot"` or `"nmpz"`.
96
- max_steps (`int`, *optional*, defaults to `24`):
97
- Actions allowed before the episode is cut off. Twelve is tight
98
- for an agentic episode: looking in four directions and walking
99
- a block spends most of it before any reasoning about the map.
100
- reward_mode (`str`, *optional*, defaults to `"coords"`):
101
- `"coords"` scores distance; `"country_only"` scores the country.
102
- max_free_calls (`int`, *optional*, defaults to `24`):
103
- How many free-tool calls (`measure`) an episode may make before they
104
- begin consuming the step budget. Free tools return arithmetic on
105
- coordinates the agent supplied and so reveal nothing, but without a
106
- cap a policy can issue them indefinitely and never terminate.
107
- reward_shape (`str`, *optional*, defaults to `"geoguessr"`):
108
- Distance curve. `"geoguessr"` is the game's own; `"mixture"` adds a
109
- 5000 km scale so a wrong-continent guess still has a gradient. Use
110
- `"mixture"` for training and `"geoguessr"` for anything you report.
111
- cost_mode (`str`, *optional*, defaults to `"subtract"`):
112
- `"subtract"` takes the action cost off the score and floors at zero,
113
- as the game does. `"multiply"` scales the score by `1 - cost`, which
114
- preserves the ordering of bad guesses -- required for training, see
115
- [`~scoring.compute_reward`].
116
- hide_task_identity (`bool`, *optional*, defaults to `False`):
117
- Drop `task_index`, `task_id`, `sequence_id` and `attribution` from
118
- the per-turn observation metadata. **Set this for RL training**: the
119
- contributor username alone determines the country for 74% of
120
- training tasks, so leaving it in lets a policy score without looking
121
- at the image. The terminal observation carries them either way.
122
- hierarchical_reward (`bool`, *optional*, defaults to `False`):
123
- Add country and region partial credit to the distance score.
124
- view_size (`int`, *optional*, defaults to `640`):
125
- Edge length in pixels of rendered views.
126
- allow_fetch (`bool`, *optional*, defaults to `True`):
127
- Whether a cache miss may reach the Mapillary API.
128
- hires_zoom (`bool`, *optional*, defaults to `True`):
129
- Render zoomed views from the full-resolution original, so a narrow
130
- field of view actually resolves detail such as distant signage.
131
- reveal_map (`bool`, *optional*, defaults to `True`):
132
- Draw a map of the guess against the truth on the terminal
133
- observation. Costs about 280 ms, which is the single largest cost in
134
- an episode, and a training run does not read it — the reward and the
135
- distance are in the observation either way. Leave it on for evals,
136
- demos and traces; turn it off for throughput.
137
-
138
- Examples:
139
-
140
- ```python
141
- env = GeoGuesserEnvironment("tasks/pano_v1.jsonl", "data/panos")
142
- observation = env.reset(task_index=0)
143
- result = env.step(LookAction(heading_deg=90))
144
- ```
145
- """
146
-
147
- SUPPORTS_CONCURRENT_SESSIONS = True
148
-
149
- def __init__(
150
- self,
151
- index_path: str | None = None,
152
- cache_dir: str = "",
153
- episode_mode: str = EpisodeMode.AGENTIC.value,
154
- max_steps: int = 24,
155
- max_free_calls: int = 24,
156
- reward_mode: str = RewardMode.COORDS.value,
157
- hierarchical_reward: bool = False,
158
- reward_shape: str = "geoguessr",
159
- cost_mode: str = "subtract",
160
- hide_task_identity: bool = False,
161
- view_size: int = 640,
162
- allow_fetch: bool = True,
163
- hires_zoom: bool = True,
164
- reveal_map: bool = True,
165
- splits: dict[str, str] | None = None,
166
- default_split: str = "train",
167
- ):
168
- if splits:
169
- self._split_paths = {name: str(path) for name, path in splits.items()}
170
- elif index_path:
171
- # Single-index callers keep working: one nameless index becomes the
172
- # default split, so existing tests, harness runs and the play UI
173
- # need no change.
174
- self._split_paths = {default_split: str(index_path)}
175
- else:
176
- raise ValueError("Provide either splits= or index_path=.")
177
- if default_split not in self._split_paths:
178
- raise ValueError(
179
- f"default_split {default_split!r} is not one of "
180
- f"{sorted(self._split_paths)}."
181
- )
182
- self._default_split = default_split
183
- self._cache_dir = cache_dir
184
- self._allow_fetch = allow_fetch
185
- self._hires_zoom = hires_zoom
186
- self._backends: dict[str, PanoramaBackend] = {}
187
- self._split = default_split
188
- self._backend = self._backend_for(default_split)
189
- self._reveal_map = reveal_map
190
- self._mode = EpisodeMode(episode_mode)
191
- self._max_steps = max_steps
192
- # Free tools reveal nothing, so they stay free -- but not unlimited.
193
- self._max_free_calls = max_free_calls
194
- self._reward_mode = RewardMode(reward_mode)
195
- self._hierarchical = hierarchical_reward
196
- self._reward_shape = reward_shape
197
- self._cost_mode = cost_mode
198
- self._hide_task_identity = hide_task_identity
199
- # Validate now rather than at the end of the first episode.
200
- compute_reward(1.0, shape=reward_shape, cost_mode=cost_mode)
201
- self._view_size = (view_size, view_size)
202
- self._state = GeoGuesserState()
203
- self._task = None
204
- self._rng = random.Random()
205
-
206
- mcp = FastMCP("geoguesser_env")
207
- self._register_tools(mcp)
208
- super().__init__(mcp)
209
-
210
- # -- splits ------------------------------------------------------------
211
-
212
- def _backend_for(self, split: str) -> PanoramaBackend:
213
- """
214
- Return the backend serving one split, building it on first use.
215
-
216
- Splits are built lazily so a deployment that only mounts the eval index
217
- is not forced to carry a training one, and because the parsed index is
218
- cached process-wide anyway.
219
-
220
- Args:
221
- split (`str`):
222
- Split name.
223
-
224
- Returns:
225
- [`PanoramaBackend`]: Backend for that split.
226
-
227
- Raises:
228
- UnknownSplitError: When the split is not configured.
229
- """
230
- if split not in self._split_paths:
231
- raise UnknownSplitError(
232
- f"Unknown split {split!r}. Available: {sorted(self._split_paths)}."
233
- )
234
- backend = self._backends.get(split)
235
- if backend is None:
236
- backend = PanoramaBackend(
237
- self._split_paths[split],
238
- self._cache_dir,
239
- allow_fetch=self._allow_fetch,
240
- hires_zoom=self._hires_zoom,
241
- )
242
- self._backends[split] = backend
243
- return backend
244
-
245
- @staticmethod
246
- def _split_type(split: str) -> str:
247
- """
248
- Map a split name onto the type vocabulary the core Task API knows.
249
-
250
- Core normalises anything outside `{train, validation, test}` to
251
- `validation`, so `eval` is declared as `test` explicitly rather than
252
- being silently downgraded.
253
- """
254
- if split == "train":
255
- return "train"
256
- if split in {"eval", "test"}:
257
- return "test"
258
- return "validation"
259
-
260
- def _task_spec(self, split: str, task: Task) -> dict[str, Any]:
261
- """
262
- Describe one task for the Task API.
263
-
264
- Deliberately truth-free: no coordinates and no country. Task specs
265
- travel to whatever orchestrates training, and the moment a label sits
266
- in a spec someone can build a prompt from it. The true location is
267
- revealed in the observation metadata after the guess, which is the one
268
- place it belongs.
269
- """
270
- return {
271
- "task_index": task.task_index,
272
- "task_id": task.task_id,
273
- "split": split,
274
- "n_frames": len(task.frames),
275
- "provider": task.provider,
276
- "sequence_id": task.sequence_id,
277
- "offline_ready": bool(task.meta.get("offline_ready", False)),
278
- }
279
-
280
- def list_splits(self) -> list[dict[str, Any]]:
281
- """
282
- Task API: describe every configured split.
283
-
284
- Returns:
285
- `list[dict]` with keys:
286
- - `name` (`str`):
287
- Split name, as accepted by `reset(split=)`.
288
- - `type` (`str`):
289
- One of `train`, `test` or `validation`.
290
- - `num_tasks` (`int`):
291
- Task count in the split.
292
- - `default` (`bool`):
293
- Whether `reset()` uses this split when none is given.
294
- """
295
- return [
296
- {
297
- "name": name,
298
- "type": self._split_type(name),
299
- "num_tasks": self._backend_for(name).n_tasks,
300
- "default": name == self._default_split,
301
- }
302
- for name in self._split_paths
303
- ]
304
-
305
- def num_tasks(self, split: str) -> int:
306
- """Task API: how many tasks a split holds."""
307
- return self._backend_for(split).n_tasks
308
-
309
- def get_task(self, split: str, index: int) -> dict[str, Any]:
310
- """Task API: describe one task by split and index."""
311
- return self._task_spec(split, self._backend_for(split).task(index))
312
-
313
- def list_tasks(self, split: str) -> list[dict[str, Any]]:
314
- """Task API: describe every task in a split."""
315
- backend = self._backend_for(split)
316
- return [
317
- self._task_spec(split, backend.task(index))
318
- for index in range(backend.n_tasks)
319
- ]
320
-
321
- def get_task_range(
322
- self, split: str, start: int | None = None, stop: int | None = None
323
- ) -> list[dict[str, Any]]:
324
- """Task API: describe a slice-style range of tasks in a split."""
325
- backend = self._backend_for(split)
326
- indices = range(*slice(start, stop).indices(backend.n_tasks))
327
- return [self._task_spec(split, backend.task(index)) for index in indices]
328
-
329
- # -- capability-aware tool registration --------------------------------
330
-
331
- def _navigational(self) -> bool:
332
- return self._mode is EpisodeMode.AGENTIC and self._backend.supports_move
333
-
334
- def _can_look(self) -> bool:
335
- return self._mode is EpisodeMode.AGENTIC and self._backend.supports_look
336
-
337
- def _can_gather(self) -> bool:
338
- """Whether the episode has an investigation phase at all.
339
-
340
- `single_shot` deliberately has none: one view, one guess, which is the
341
- shape a VLM GRPO run wants. `nmpz` keeps the map but takes the camera
342
- away, mirroring the game's own hardest mode.
343
- """
344
- return self._mode is not EpisodeMode.SINGLE_SHOT
345
-
346
- def _register_tools(self, mcp: FastMCP) -> None:
347
- """Register only the tools this configuration can actually serve."""
348
- if self._can_look():
349
-
350
- @mcp.tool
351
- def look(
352
- heading_deg: float, pitch_deg: float = 0.0, fov_deg: float = 90.0
353
- ) -> str:
354
- """Look in a direction. heading_deg is absolute, 0 = true north.
355
-
356
- Args:
357
- heading_deg: Compass heading in degrees.
358
- pitch_deg: Vertical angle; positive looks up.
359
- fov_deg: Field of view; smaller values zoom in.
360
- """
361
- return self._apply(
362
- LookAction(
363
- heading_deg=heading_deg, pitch_deg=pitch_deg, fov_deg=fov_deg
364
- )
365
- ).feedback
366
-
367
- @mcp.tool
368
- def pan(delta_deg: float) -> str:
369
- """Turn relative to the current heading; positive turns right.
370
-
371
- Args:
372
- delta_deg: Degrees to turn.
373
- """
374
- return self._apply(PanAction(delta_deg=delta_deg)).feedback
375
-
376
- @mcp.tool
377
- def zoom(fov_deg: float) -> str:
378
- """Change field of view without turning. 30 reads distant signs.
379
-
380
- Args:
381
- fov_deg: New field of view in degrees.
382
- """
383
- return self._apply(ZoomAction(fov_deg=fov_deg)).feedback
384
-
385
- if self._navigational():
386
-
387
- @mcp.tool
388
- def move(direction: str, meters: float = 10.0) -> str:
389
- """Walk along the road. Reports how far you actually travelled.
390
-
391
- Args:
392
- direction: Either "forward" or "backward".
393
- meters: Requested distance in metres.
394
- """
395
- return self._apply(
396
- MoveAction(direction=direction, meters=meters)
397
- ).feedback
398
-
399
- if self._can_gather():
400
- self._register_map_tools(mcp)
401
- self._register_guess_tool(mcp)
402
-
403
- def _register_map_tools(self, mcp: FastMCP) -> None:
404
- """Register the map and pin tools, which every gathering mode has."""
405
-
406
- @mcp.tool
407
- def place_pin(
408
- lat: float, lon: float, label: str = "", span_deg: float = 7.0
409
- ) -> str:
410
- """Pin a candidate and see where it falls on the map.
411
-
412
- Tells you what is at that coordinate. Says nothing about whether
413
- you are right.
414
-
415
- Args:
416
- lat: Latitude of the candidate.
417
- lon: Longitude of the candidate.
418
- label: Optional note.
419
- span_deg: Half-width of the map window in degrees. Below about
420
- 4 the map adds roads, urban areas and town names, which is
421
- how you aim within a city rather than at its centre.
422
- """
423
- return self._apply(
424
- PinAction(lat=lat, lon=lon, label=label or None, span_deg=span_deg)
425
- ).feedback
426
-
427
- @mcp.tool
428
- def view_map(lat: float, lon: float, span_deg: float = 7.0) -> str:
429
- """Pan and zoom the map without placing a pin.
430
-
431
- Args:
432
- lat: Latitude at the centre of the view.
433
- lon: Longitude at the centre of the view.
434
- span_deg: Half-width of the window in degrees.
435
- """
436
- return self._apply(
437
- ViewMapAction(lat=lat, lon=lon, span_deg=span_deg)
438
- ).feedback
439
-
440
- @mcp.tool
441
- def list_pins() -> str:
442
- """List the candidates pinned so far. Free."""
443
- if not self._state.pins:
444
- return "No pins placed yet."
445
- return "\n".join(
446
- f"{i}. {p['lat']:.4f}, {p['lon']:.4f} - {p['description']}"
447
- for i, p in enumerate(self._state.pins, 1)
448
- )
449
-
450
- @mcp.tool
451
- def clear_pins() -> str:
452
- """Remove all pins. Free."""
453
- self._state.pins = []
454
- return "Pins cleared."
455
-
456
- @mcp.tool
457
- def measure(lat_a: float, lon_a: float, lat_b: float, lon_b: float) -> str:
458
- """Distance in km between two coordinates of your own choosing. Free.
459
-
460
- Args:
461
- lat_a: Latitude of the first point.
462
- lon_a: Longitude of the first point.
463
- lat_b: Latitude of the second point.
464
- lon_b: Longitude of the second point.
465
- """
466
- km = haversine_km(lat_a, lon_a, lat_b, lon_b)
467
- return f"{km:.0f} km between those two points."
468
-
469
- @mcp.tool
470
- def reverse_geocode(lat: float, lon: float) -> str:
471
- """Name the country and nearest city at a coordinate. Free.
472
-
473
- Args:
474
- lat: Latitude in degrees.
475
- lon: Longitude in degrees.
476
- """
477
- place = locate(lat, lon)
478
- where = place.country or "open water"
479
- return (
480
- f"{lat:.4f}, {lon:.4f} is in {where}. Nearest major city: "
481
- f"{place.nearest_city}, ~{place.city_distance_km:.0f} km "
482
- f"{place.city_bearing}."
483
- )
484
-
485
- def _register_guess_tool(self, mcp: FastMCP) -> None:
486
- """Register the terminal action, which every mode has."""
487
-
488
- @mcp.tool
489
- def submit_guess(
490
- lat: float,
491
- lon: float,
492
- country: str = "",
493
- confidence: float = -1.0,
494
- reasoning: str = "",
495
- ) -> str:
496
- """Commit your final answer. Ends the episode.
497
-
498
- Args:
499
- lat: Latitude of your guess.
500
- lon: Longitude of your guess.
501
- country: Optional ISO-3166 alpha-2 code or country name.
502
- confidence: Optional self-reported confidence in [0, 1].
503
- reasoning: Optional rationale, recorded but not scored.
504
- """
505
- observation = self._apply(
506
- GuessAction(
507
- lat=lat,
508
- lon=lon,
509
- country=country or None,
510
- confidence=None if confidence < 0 else confidence,
511
- reasoning=reasoning or None,
512
- )
513
- )
514
- return observation.feedback
515
-
516
- # -- lifecycle ---------------------------------------------------------
517
-
518
- def reset(
519
- self,
520
- seed: int | None = None,
521
- episode_id: str | None = None,
522
- split: str | None = None,
523
- index: int | None = None,
524
- task_index: int | None = None,
525
- **kwargs: Any,
526
- ) -> GeoGuesserObservation:
527
- """
528
- Start an episode.
529
-
530
- Selection is explicit, because training and demoing want opposite
531
- things. `index` picks one exact task and is byte-identical on repeat,
532
- which is what a GRPO group needs. `seed` picks `tasks[seed % n_tasks]`.
533
- Neither means a random task; omitting both does, and the chosen split
534
- and index are always reported in the observation metadata so a random
535
- episode stays replayable.
536
-
537
- Args:
538
- seed (`int`, *optional*):
539
- Deterministic selector, `tasks[seed % n_tasks]`.
540
- episode_id (`str`, *optional*):
541
- Caller-supplied episode identifier.
542
- split (`str`, *optional*):
543
- Which split to draw from. Defaults to the environment's default
544
- split.
545
- index (`int`, *optional*):
546
- Exact task to play, within `split`. Takes precedence over
547
- `seed`.
548
- task_index (`int`, *optional*):
549
- Deprecated alias for `index`, kept so existing callers and
550
- saved trajectories keep working.
551
-
552
- Returns:
553
- [`GeoGuesserObservation`]: The opening view and prompt.
554
-
555
- Raises:
556
- UnknownSplitError: When `split` is not configured.
557
- """
558
- self._split = split or self._default_split
559
- self._backend = self._backend_for(self._split)
560
-
561
- if index is None:
562
- index = task_index
563
- n = self._backend.n_tasks
564
- if index is not None:
565
- chosen = int(index) % n
566
- elif seed is not None:
567
- chosen = int(seed) % n
568
- else:
569
- chosen = self._rng.randrange(n)
570
-
571
- self._task = self._backend.task(chosen)
572
- start = self._task.frames[self._task.start_frame]
573
- self._state = GeoGuesserState(
574
- episode_id=episode_id or str(uuid.uuid4()),
575
- task_index=chosen,
576
- task_id=self._task.task_id,
577
- frame_index=self._task.start_frame,
578
- heading_deg=0.0,
579
- pitch_deg=0.0,
580
- fov_deg=90.0,
581
- )
582
-
583
- observation = self._render_view_observation()
584
- observation.prompt = PROMPT.format(
585
- tools=", ".join(self._tool_names()), steps=self._max_steps
586
- )
587
- observation.feedback = "Episode started."
588
- observation.captured_at = start.captured_at
589
- observation.metadata = self._metadata()
590
- return observation
591
-
592
- @property
593
- def state(self) -> GeoGuesserState:
594
- """Current internal state."""
595
- return self._state
596
-
597
- def _step_impl(self, action: Action, **kwargs: Any) -> Observation:
598
- """Handle structured, non-MCP actions.
599
-
600
- Accepts either the flat wire action the HTTP layer delivers or a typed
601
- action constructed in process, so tests and the harness can bypass
602
- serialisation without a second code path.
603
- """
604
- if isinstance(action, GeoGuesserAction):
605
- return self._apply(from_wire(action))
606
- if isinstance(action, TypedAction):
607
- return self._apply(action)
608
- raise TypeError(f"Unsupported action type: {type(action).__name__}")
609
-
610
- # -- the actual mechanics ---------------------------------------------
611
-
612
- def _tool_names(self) -> list[str]:
613
- """Names of the tools this configuration actually registered."""
614
- names: list[str] = []
615
- if self._can_look():
616
- names += ["look", "pan", "zoom"]
617
- if self._navigational():
618
- names.append("move")
619
- if self._can_gather():
620
- names += [
621
- "place_pin",
622
- "view_map",
623
- "list_pins",
624
- "clear_pins",
625
- "measure",
626
- "reverse_geocode",
627
- ]
628
- names.append("submit_guess")
629
- return names
630
-
631
- def _steps_used(self) -> int:
632
- s = self._state
633
- return s.n_looks + s.n_maps + s.n_pins + s.n_moves
634
-
635
- def _steps_remaining(self) -> int:
636
- return max(0, self._max_steps - self._steps_used())
637
-
638
- def _cost(self) -> float:
639
- s = self._state
640
- return action_cost(
641
- n_looks=s.n_looks, n_maps=s.n_maps, n_pins=s.n_pins, n_moves=s.n_moves
642
- )
643
-
644
- def _metadata(self) -> dict[str, Any]:
645
- # Identity fields are a reward-hacking channel, not just clutter. The
646
- # Mapillary contributor determines the country outright for 74% of
647
- # training tasks ("amsterdam" only maps the Netherlands), and
648
- # task_index/task_id/sequence_id are a few thousand memorisable keys
649
- # straight to a coordinate -- either lets a policy score without ever
650
- # reading the image. Kept by default so the play UI can show the licence
651
- # credit and recordings keep their provenance; RL training must set
652
- # hide_task_identity=True. The terminal observation carries them
653
- # regardless, since by then the truth is already revealed.
654
- if self._hide_task_identity:
655
- return {
656
- "split": self._split,
657
- "street_detail": (
658
- "unavailable"
659
- if street_fetch_failed()
660
- else ("on" if street_detail_enabled() else "off")
661
- ),
662
- "backend": "mapillary",
663
- "episode_mode": self._mode.value,
664
- "frame_index": self._state.frame_index,
665
- }
666
- return {
667
- "split": self._split,
668
- # A map quietly missing its streets looks like a styling choice, so
669
- # say so. Overpass 504s from datacenter egress, which is how a Space
670
- # ends up with poorer maps than a laptop for the same task.
671
- "street_detail": (
672
- "unavailable"
673
- if street_fetch_failed()
674
- else ("on" if street_detail_enabled() else "off")
675
- ),
676
- "task_index": self._state.task_index,
677
- "task_id": self._state.task_id,
678
- "backend": "mapillary",
679
- # Opaque upstream identifier, not ground truth: it makes a recorded
680
- # episode traceable back to its source sequence.
681
- "sequence_id": self._task.sequence_id if self._task else None,
682
- "episode_mode": self._mode.value,
683
- "frame_index": self._state.frame_index,
684
- "captured_at": self._task.frames[self._state.frame_index].captured_at,
685
- "attribution": self._task.attribution,
686
- }
687
-
688
- def _base_observation(self) -> GeoGuesserObservation:
689
- s = self._state
690
- return GeoGuesserObservation(
691
- heading_deg=s.heading_deg % 360,
692
- pitch_deg=s.pitch_deg,
693
- fov_deg=s.fov_deg,
694
- total_moved_meters=s.total_moved_meters,
695
- can_move_forward=self._navigational()
696
- and self._backend.can_move(self._task, s.frame_index, "forward"),
697
- can_move_backward=self._navigational()
698
- and self._backend.can_move(self._task, s.frame_index, "backward"),
699
- available_tools=self._tool_names(),
700
- steps_remaining=self._steps_remaining(),
701
- action_cost=self._cost(),
702
- pins=[Pin(**p) for p in s.pins],
703
- captured_at=self._task.frames[s.frame_index].captured_at,
704
- metadata=self._metadata(),
705
- )
706
-
707
- def _render_view_observation(self) -> GeoGuesserObservation:
708
- s = self._state
709
- view = self._backend.render_view(
710
- self._task, s.frame_index, s.heading_deg, s.pitch_deg, s.fov_deg
711
- )
712
- if view.size != self._view_size:
713
- view = view.resize(self._view_size)
714
- observation = self._base_observation()
715
- observation.image_base64 = to_base64(view, "JPEG")
716
- observation.image_kind = "view"
717
- return observation
718
-
719
- def _render_map_observation(
720
- self,
721
- pins: list[tuple[float, float]],
722
- focus,
723
- span: float,
724
- truth: tuple[float, float] | None = None,
725
- ) -> GeoGuesserObservation:
726
- image = render_map(pins, focus=focus, span_deg=span, truth=truth)
727
- observation = self._base_observation()
728
- observation.image_base64 = to_base64(image, "PNG")
729
- observation.image_kind = "map"
730
- return observation
731
-
732
- def _apply(self, action: Action) -> GeoGuesserObservation:
733
- """Execute one action and produce the resulting observation."""
734
- if self._task is None:
735
- raise RuntimeError("reset() must be called before step().")
736
-
737
- s = self._state
738
- s.step_count += 1
739
-
740
- if s.submitted:
741
- observation = self._base_observation()
742
- observation.done = True
743
- observation.feedback = "The episode is over; the guess was already made."
744
- return observation
745
-
746
- if isinstance(action, GuessAction):
747
- return self._finish(action)
748
-
749
- if self._steps_remaining() <= 0:
750
- observation = self._base_observation()
751
- observation.feedback = (
752
- "Out of actions. Call submit_guess with your best estimate."
753
- )
754
- return observation
755
-
756
- if isinstance(action, LookAction):
757
- s.heading_deg = action.heading_deg
758
- s.pitch_deg = action.pitch_deg
759
- s.fov_deg = action.fov_deg
760
- s.n_looks += 1
761
- observation = self._render_view_observation()
762
- observation.feedback = (
763
- f"Facing {s.heading_deg % 360:.0f} deg, {s.fov_deg:.0f} deg field of "
764
- f"view. {self._steps_remaining()} actions left."
765
- )
766
- return observation
767
-
768
- if isinstance(action, PanAction):
769
- s.heading_deg = (s.heading_deg + action.delta_deg) % 360
770
- s.n_looks += 1
771
- observation = self._render_view_observation()
772
- observation.feedback = (
773
- f"Turned to {s.heading_deg:.0f} deg. "
774
- f"{self._steps_remaining()} actions left."
775
- )
776
- return observation
777
-
778
- if isinstance(action, ZoomAction):
779
- s.fov_deg = action.fov_deg
780
- s.n_looks += 1
781
- observation = self._render_view_observation()
782
- observation.feedback = (
783
- f"Field of view now {s.fov_deg:.0f} deg. "
784
- f"{self._steps_remaining()} actions left."
785
- )
786
- return observation
787
-
788
- if isinstance(action, MoveAction):
789
- new_index, travelled = self._backend.step_along(
790
- self._task, s.frame_index, action.direction, action.meters
791
- )
792
- s.n_moves += 1
793
- if new_index == s.frame_index:
794
- observation = self._render_view_observation()
795
- observation.feedback = (
796
- f"Cannot go {action.direction} - the captured road ends here. "
797
- f"{self._steps_remaining()} actions left."
798
- )
799
- return observation
800
- s.frame_index = new_index
801
- s.total_moved_meters += travelled
802
- observation = self._render_view_observation()
803
- observation.moved_meters = travelled
804
- observation.feedback = (
805
- f"Moved {travelled:.0f} m {action.direction} "
806
- f"({s.total_moved_meters:.0f} m total). "
807
- f"{self._steps_remaining()} actions left."
808
- )
809
- return observation
810
-
811
- if isinstance(action, PinAction):
812
- previous = (s.pins[-1]["lat"], s.pins[-1]["lon"]) if s.pins else None
813
- description = describe_pin(
814
- len(s.pins) + 1, action.lat, action.lon, previous=previous
815
- )
816
- s.pins.append(
817
- {
818
- "index": len(s.pins) + 1,
819
- "lat": action.lat,
820
- "lon": action.lon,
821
- "label": action.label,
822
- "description": description,
823
- }
824
- )
825
- s.n_pins += 1
826
- pins = [(p["lat"], p["lon"]) for p in s.pins]
827
- observation = self._render_map_observation(
828
- pins, (action.lat, action.lon), action.span_deg
829
- )
830
- observation.feedback = (
831
- f"{description} {self._steps_remaining()} actions left."
832
- )
833
- return observation
834
-
835
- if isinstance(action, ViewMapAction):
836
- s.n_maps += 1
837
- pins = [(p["lat"], p["lon"]) for p in s.pins]
838
- observation = self._render_map_observation(
839
- pins, (action.lat, action.lon), action.span_deg
840
- )
841
- place = locate(action.lat, action.lon)
842
- observation.feedback = (
843
- f"Map centred on {action.lat:.3f}, {action.lon:.3f} "
844
- f"({place.country or 'open water'}), "
845
- f"{action.span_deg * 2:.0f} deg across. "
846
- f"{self._steps_remaining()} actions left."
847
- )
848
- return observation
849
-
850
- if isinstance(action, MeasureAction):
851
- # Free, because it is arithmetic on two coordinates the agent
852
- # supplied: it reveals nothing about where the agent is. But free
853
- # used to mean unbounded -- it incremented no counter, so it never
854
- # advanced the step budget and a policy could issue it forever. Past
855
- # the cap it starts costing a map action, so the episode terminates.
856
- if s.n_free >= self._max_free_calls:
857
- s.n_maps += 1
858
- else:
859
- s.n_free += 1
860
- km = haversine_km(action.lat_a, action.lon_a, action.lat_b, action.lon_b)
861
- observation = self._base_observation()
862
- observation.feedback = f"{km:.0f} km between those two points."
863
- if s.n_free >= self._max_free_calls:
864
- observation.feedback += " Free-tool budget spent; this now costs."
865
- return observation
866
-
867
- raise TypeError(f"Unsupported action type: {type(action).__name__}")
868
-
869
- def _finish(self, action: GuessAction) -> GeoGuesserObservation:
870
- """Score the final guess and end the episode."""
871
- s = self._state
872
- s.submitted = True
873
- true_lat, true_lon = self._task.truth
874
-
875
- if action.lat is not None and action.lon is not None:
876
- lat, lon, parsed_ok, note = action.lat, action.lon, True, ""
877
- else:
878
- parsed = parse_guess(action.response or "")
879
- lat, lon, parsed_ok, note = parsed.lat, parsed.lon, parsed.ok, parsed.note
880
-
881
- cost = self._cost()
882
- observation = self._base_observation()
883
- observation.done = True
884
- observation.parsed_ok = parsed_ok
885
- observation.true_lat = true_lat
886
- observation.true_lon = true_lon
887
- observation.action_cost = cost
888
-
889
- if not parsed_ok:
890
- observation.reward = 0.0
891
- observation.score = 0.0
892
- observation.feedback = f"No usable guess. {note} Scored 0."
893
- observation.metadata = {
894
- **self._metadata(),
895
- "parse_failure": True,
896
- "country": self._task.country,
897
- "task_index": self._state.task_index,
898
- "task_id": self._state.task_id,
899
- "sequence_id": self._task.sequence_id,
900
- "attribution": self._task.attribution,
901
- }
902
- return observation
903
-
904
- distance = haversine_km(lat, lon, true_lat, true_lon)
905
-
906
- # A guess previously returned no image at all, which left the outcome
907
- # invisible in a trace and gave a policy nothing to learn the shape of
908
- # its error from. Truth is only ever drawn here, after scoring.
909
- if not self._reveal_map:
910
- observation.image_kind = "none"
911
- separation = max(abs(lat - true_lat), abs(lon - true_lon))
912
- if self._reveal_map and separation > 25.0:
913
- # Framing both points would squash a hemisphere into the panel and
914
- # tell you nothing. The useful second view is where it actually was.
915
- reveal_focus = (true_lat, true_lon)
916
- reveal_span = 12.0
917
- elif self._reveal_map:
918
- reveal_focus = ((lat + true_lat) / 2, (lon + true_lon) / 2)
919
- reveal_span = max(0.05, separation * 0.75 + 0.4)
920
- if self._reveal_map:
921
- reveal = self._render_map_observation(
922
- [(lat, lon)], reveal_focus, reveal_span, truth=(true_lat, true_lon)
923
- )
924
- observation.image_base64 = reveal.image_base64
925
- observation.image_kind = "map"
926
-
927
- truth_place = locate(true_lat, true_lon)
928
- guess_place = locate(lat, lon)
929
- country_hit = bool(
930
- truth_place.country and truth_place.country == guess_place.country
931
- )
932
- region_hit = bool(
933
- truth_place.subregion and truth_place.subregion == guess_place.subregion
934
- )
935
-
936
- if self._reward_mode is RewardMode.COUNTRY_ONLY:
937
- reward = max(0.0, float(country_hit) - cost)
938
- score = float(country_hit)
939
- else:
940
- score = compute_reward(
941
- distance,
942
- cost=0.0,
943
- country_hit=country_hit,
944
- region_hit=region_hit,
945
- hierarchical=self._hierarchical,
946
- shape=self._reward_shape,
947
- )
948
- reward = compute_reward(
949
- distance,
950
- cost=cost,
951
- country_hit=country_hit,
952
- region_hit=region_hit,
953
- hierarchical=self._hierarchical,
954
- shape=self._reward_shape,
955
- cost_mode=self._cost_mode,
956
- )
957
-
958
- observation.distance_km = distance
959
- observation.score = score
960
- observation.reward = reward
961
- observation.feedback = (
962
- f"{verdict(distance)} - {distance:.0f} km away. True location "
963
- f"{true_lat:.4f}, {true_lon:.4f} "
964
- f"({truth_place.country or 'open water'}). "
965
- + (
966
- f"Score {score:.3f} scaled by {1 - min(cost, 0.5):.2f} "
967
- f"action cost = {reward:.3f}."
968
- if self._cost_mode == "multiply"
969
- else f"Score {score:.3f} minus {cost:.2f} action cost = {reward:.3f}."
970
- )
971
- )
972
- observation.metadata = {
973
- **self._metadata(),
974
- # Revealed here and only here, next to true_lat/true_lon: the
975
- # country is the label a per-region breakdown needs, and a task spec
976
- # deliberately does not carry it.
977
- "country": self._task.country,
978
- # Restored here even under hide_task_identity: the episode is over,
979
- # so provenance can no longer be used to shortcut it.
980
- "task_index": self._state.task_index,
981
- "task_id": self._state.task_id,
982
- "sequence_id": self._task.sequence_id,
983
- "attribution": self._task.attribution,
984
- "guess": [lat, lon],
985
- "guess_country": guess_place.country,
986
- "verdict": verdict(distance),
987
- "distance_km": distance,
988
- "country_hit": country_hit,
989
- "region_hit": region_hit,
990
- "confidence": action.confidence,
991
- "reasoning": action.reasoning,
992
- "n_looks": s.n_looks,
993
- "n_maps": s.n_maps,
994
- "n_pins": s.n_pins,
995
- "n_moves": s.n_moves,
996
- "total_moved_meters": s.total_moved_meters,
997
- }
998
- return observation
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/gradio_ui.py DELETED
@@ -1,1044 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """The human-play page: a GeoGuessr-style game in the browser.
4
-
5
- Two viewers, both driven by the environment's own data:
6
-
7
- - Pannellum shows the equirectangular panorama, so a person drags to look
8
- around exactly where the agent calls `look()`.
9
- - MapLibre shows the guess map over OpenFreeMap tiles. No API key, no request
10
- limits, commercial use permitted, and self-hostable if the public instance
11
- ever goes away.
12
-
13
- The page is served as its own document at `/geoguesser/play` and embedded in
14
- the Gradio tab through an iframe, because `gr.HTML` inserts markup without
15
- executing `<script>` tags — styles apply, but neither viewer initialises, which
16
- looks like a blank panel and reports no error anywhere.
17
-
18
- A human sees live tiles; the agent's map stays the offline Natural Earth
19
- render. They agree on geometry, which is what matters, and the agent keeps a
20
- determinism the browser does not need.
21
- """
22
-
23
- from __future__ import annotations
24
-
25
- import json
26
- import random
27
- import urllib.parse
28
- from typing import Any, Dict, List, Optional
29
-
30
- import gradio as gr
31
-
32
-
33
- MAPLIBRE_JS = "https://cdnjs.cloudflare.com/ajax/libs/maplibre-gl/5.24.0/maplibre-gl.js"
34
- MAPLIBRE_CSS = (
35
- "https://cdnjs.cloudflare.com/ajax/libs/maplibre-gl/5.24.0/maplibre-gl.css"
36
- )
37
- PANNELLUM_JS = "https://cdnjs.cloudflare.com/ajax/libs/pannellum/2.5.6/pannellum.js"
38
- PANNELLUM_CSS = "https://cdnjs.cloudflare.com/ajax/libs/pannellum/2.5.6/pannellum.css"
39
- OPENFREEMAP_STYLE = "https://tiles.openfreemap.org/styles/positron"
40
-
41
- # The real game scores a round out of 5000 on this curve. Showing points
42
- # rather than the RL reward makes a score comparable to GeoGuessr intuition;
43
- # the reward is shown beside it so the two are never confused. There is no
44
- # multi-round game here, because an episode is exactly one guess.
45
- MAX_POINTS_PER_ROUND = 5000
46
-
47
- _TEMPLATE = r"""<!doctype html>
48
- <html lang="en">
49
- <head>
50
- <meta charset="utf-8">
51
- <meta name="viewport" content="width=device-width,initial-scale=1">
52
- <title>geoguesser_env - play</title>
53
- <link rel="stylesheet" href="__MAPLIBRE_CSS__">
54
- <link rel="stylesheet" href="__PANNELLUM_CSS__">
55
- <style>
56
- :root {
57
- --panel: rgba(18, 23, 28, .92);
58
- --edge: #2c353d;
59
- --ink: #e8ecef;
60
- --ink-soft: #9aa5ad;
61
- --ink-faint: #6e7a83;
62
- --pin: #c4332a;
63
- --good: #5aa06e;
64
- --wire: #7fb3cc;
65
- }
66
- * { box-sizing: border-box; }
67
- html, body {
68
- margin: 0; height: 100%; overflow: hidden; background: #10151a;
69
- color: var(--ink);
70
- font-family: ui-monospace, "SF Mono", "IBM Plex Mono", Menlo, monospace;
71
- }
72
- #stage { position: absolute; inset: 0; }
73
- #pano { position: absolute; inset: 0; }
74
- .pnlm-zoom-controls, .pnlm-orientation-button, .pnlm-panorama-info,
75
- .pnlm-compass { display: none !important; }
76
- .pnlm-load-box { background: #10151a !important; }
77
-
78
- .hud {
79
- position: absolute; z-index: 5; background: var(--panel);
80
- border: 1px solid var(--edge); border-radius: 4px;
81
- font-size: 12px; padding: 7px 11px; backdrop-filter: blur(8px);
82
- line-height: 1.5;
83
- }
84
- .hud b { color: #fff; font-weight: 500; }
85
- .hud span { color: var(--ink-soft); }
86
- #top { top: 12px; left: 12px; }
87
- #top .ep { color: var(--wire); }
88
- #compass { top: 12px; left: 50%; transform: translateX(-50%); letter-spacing: .1em; }
89
- #score { top: 12px; right: 12px; text-align: right; }
90
- #score .pts { font-size: 15px; color: #fff; }
91
- #credit {
92
- top: 74px; right: 12px; font-size: 10.5px; color: var(--ink-soft);
93
- max-width: 34vw; text-align: right; z-index: 14;
94
- }
95
- #credit a { color: #8fb8cc; text-decoration: none; }
96
- #actions {
97
- bottom: 12px; left: 50%; transform: translateX(-50%); display: flex;
98
- gap: 6px; align-items: center; transition: opacity .3s ease;
99
- }
100
- #actions {
101
- gap: 0; padding: 0; display: flex; align-items: stretch; overflow: hidden;
102
- bottom: 12px; left: 12px; transform: none;
103
- }
104
- /* One group per kind of environment action, each labelled, each showing what
105
- it costs through its tooltip rather than shouting a number. A group whose
106
- tool is not registered is removed rather than greyed: a control you cannot
107
- use is noise. */
108
- .pad {
109
- display: flex; flex-direction: column; gap: 4px; padding: 8px 14px;
110
- justify-content: center;
111
- }
112
- .pad + .pad { border-left: 1px solid var(--edge); }
113
- .pad.gone { display: none; }
114
- .padlabel {
115
- font-size: 9.5px; letter-spacing: .14em; text-transform: uppercase;
116
- color: var(--ink-faint);
117
- }
118
- .padlabel b { color: var(--ink); font-weight: 500; letter-spacing: 0; }
119
- .btns { display: flex; align-items: center; gap: 5px; }
120
- #actions button {
121
- font-family: inherit; cursor: pointer; color: var(--ink);
122
- background: #232c33; border: 1px solid var(--edge); border-radius: 4px;
123
- display: flex; flex-direction: column; align-items: center; gap: 1px;
124
- min-width: 46px; padding: 5px 7px; line-height: 1;
125
- }
126
- #actions button .glyph { font-size: 12px; }
127
- #actions button .tag {
128
- font-size: 8.5px; letter-spacing: .06em; color: var(--ink-soft);
129
- }
130
- #actions button:hover {
131
- border-color: #6d8493; background: #2b353d; color: #fff;
132
- }
133
- #actions button:hover .tag { color: var(--ink); }
134
- #actions button:active { background: #1c242a; }
135
- #actions button:focus-visible { outline: 2px solid var(--wire); outline-offset: 1px; }
136
- #actions button.gone { display: none; }
137
- #actions.working { opacity: .55; }
138
- #actions.working button { cursor: progress; }
139
- #actions.working #padBudget b { color: var(--wire); }
140
- #actions button.preset {
141
- font-size: 10px; letter-spacing: .04em; min-width: 44px; padding: 7px 8px;
142
- }
143
- #actions button.preset.on { border-color: var(--wire); color: #fff; }
144
- #fovRange { width: 104px; accent-color: var(--wire); margin-left: 4px; }
145
- .budget { gap: 4px; font-size: 10px; color: var(--ink-faint); white-space: nowrap; }
146
- .budget .sep { color: var(--edge); margin: 0 3px; }
147
- .budget b {
148
- font-size: 14px; color: var(--ink); font-weight: 500;
149
- font-variant-numeric: tabular-nums;
150
- }
151
-
152
- /* ---- rollout trace: the observation stream the agent would receive ---- */
153
- /* The trace lives on the left and the guess map on the right, so an
154
- expanded map can never cover the observation stream. */
155
- #trace {
156
- position: absolute; top: 58px; left: 12px; width: 310px; z-index: 7;
157
- max-height: calc(100% - 130px); display: flex; flex-direction: column;
158
- background: var(--panel); border: 1px solid var(--edge); border-radius: 4px;
159
- backdrop-filter: blur(8px); overflow: hidden;
160
- }
161
- #trace.hidden { display: none; }
162
- #trace h4 {
163
- margin: 0; padding: 8px 11px; font-size: 10.5px; font-weight: 500;
164
- letter-spacing: .12em; text-transform: uppercase; color: var(--wire);
165
- border-bottom: 1px solid var(--edge); display: flex;
166
- justify-content: space-between;
167
- }
168
- #trace h4 em { color: var(--ink-faint); font-style: normal; letter-spacing: 0; }
169
- #trace h4 > span:last-child { display: flex; align-items: center; gap: 8px; }
170
- #collapse {
171
- font-family: inherit; font-size: 13px; line-height: 1; cursor: pointer;
172
- background: #232c33; border: 1px solid var(--edge); border-radius: 3px;
173
- color: var(--ink); width: 22px; height: 19px; padding: 0;
174
- }
175
- #expand {
176
- font-family: inherit; font-size: 11px; line-height: 1; cursor: pointer;
177
- background: #232c33; border: 1px solid var(--edge); border-radius: 3px;
178
- color: var(--ink); padding: 5px 9px;
179
- }
180
- #collapse:hover, #expand:hover { border-color: #5b6b78; color: #fff; }
181
- /* With the trace hidden, a small control stays where it was. */
182
- #expandWrap {
183
- position: absolute; top: 58px; left: 12px; z-index: 7; display: none;
184
- padding: 4px 5px;
185
- }
186
- #expandWrap.show { display: block; }
187
- #steps { overflow-y: auto; padding: 4px 0; font-size: 11px; }
188
- .step { padding: 6px 11px; border-bottom: 1px solid #1e262c; line-height: 1.5; }
189
- .step:last-child { border-bottom: none; }
190
- .step .op { color: var(--wire); }
191
- .step .rw { float: right; color: var(--ink-faint); }
192
- .step .fb { color: var(--ink-soft); display: block; margin-top: 2px; }
193
- #agentview { border-top: 1px solid var(--edge); padding: 8px 11px; }
194
- #agentview .cap {
195
- font-size: 9.5px; letter-spacing: .1em; text-transform: uppercase;
196
- color: var(--ink-faint); margin-bottom: 5px;
197
- }
198
- #agentview img { width: 100%; display: block; border-radius: 3px; }
199
-
200
- /* ---- guess map ------------------------------------------------------- */
201
- #mapwrap {
202
- position: absolute; right: 12px; bottom: 12px; z-index: 6;
203
- width: 300px; height: 210px; border: 1px solid var(--edge);
204
- border-radius: 5px; overflow: hidden; opacity: .9; background: #191f24;
205
- transition: width .24s ease, height .24s ease, opacity .24s ease,
206
- right .24s ease, bottom .24s ease;
207
- }
208
- #mapwrap:hover, #mapwrap.big {
209
- width: min(560px, 48vw); height: min(400px, 58vh); opacity: 1;
210
- }
211
- #mapwrap.reveal {
212
- right: 50%; bottom: 50%; transform: translate(50%, 50%);
213
- width: min(1000px, 86vw); height: min(600px, 72vh); opacity: 1;
214
- }
215
- /* The reveal is the whole point of the round, so it may cover the trace. */
216
- #mapwrap.reveal { z-index: 13; }
217
- #map { position: absolute; inset: 0; }
218
- #submit {
219
- position: absolute; left: 0; right: 0; bottom: 0; z-index: 3; width: 100%;
220
- font-family: inherit; font-size: 12px; letter-spacing: .07em;
221
- text-transform: uppercase; padding: 10px; border: none;
222
- border-top: 1px solid var(--edge); background: #2b3540; color: #7e8b96;
223
- cursor: not-allowed;
224
- }
225
- #submit.ready { background: var(--pin); color: #fff; cursor: pointer; }
226
- #submit.ready:hover { background: #d64236; }
227
- #mapwrap.reveal #submit { display: none; }
228
-
229
- /* ---- result bar ------------------------------------------------------ */
230
- #result {
231
- position: absolute; left: 0; right: 0; bottom: 0; z-index: 12; display: none;
232
- background: rgba(14, 18, 22, .96); border-top: 1px solid var(--edge);
233
- padding: 13px 20px; backdrop-filter: blur(8px);
234
- }
235
- #result.show { display: block; }
236
- #result .inner {
237
- max-width: 1180px; margin: 0 auto; display: flex; align-items: center;
238
- gap: 22px; flex-wrap: wrap;
239
- }
240
- #verdict { font-size: 17px; color: #fff; min-width: 140px; }
241
- .stat { font-size: 10.5px; color: var(--ink-soft); }
242
- .stat b {
243
- display: block; font-size: 14px; color: var(--ink); font-weight: 500;
244
- font-variant-numeric: tabular-nums; margin-top: 2px;
245
- }
246
- .stat.env b { color: var(--wire); }
247
- #bar { flex: 1 1 140px; min-width: 100px; height: 6px; background: #232b32;
248
- border-radius: 4px; overflow: hidden; }
249
- #bar i { display: block; height: 100%; background: var(--good); width: 0; }
250
- #advance {
251
- font-family: inherit; font-size: 12px; text-transform: uppercase;
252
- letter-spacing: .06em; padding: 10px 18px; background: #232b32;
253
- color: var(--ink); border: 1px solid var(--edge); border-radius: 3px;
254
- cursor: pointer;
255
- }
256
- #advance:hover { border-color: #465360; color: #fff; }
257
-
258
- .dot {
259
- display: inline-block; width: 7px; height: 7px; border-radius: 50%;
260
- margin-right: 7px; background: var(--ink-faint);
261
- vertical-align: 1px;
262
- }
263
- .dot.live { background: var(--good); }
264
- .dot.bad { background: var(--pin); }
265
- </style>
266
- </head>
267
- <body>
268
- <div id="stage">
269
- <div id="pano"></div>
270
-
271
- <div class="hud" id="top">
272
- <i id="conn" class="dot" title="environment session"></i><b>geoguesser_env</b>
273
- <span>episode</span> <b class="ep" id="task">-</b>
274
- <span>· frame</span> <b id="frame">-</b>
275
- <span>· steps</span> <b id="stepsLeft">-</b>
276
- </div>
277
- <div class="hud" id="compass"><span>facing</span> <b id="heading">-</b>
278
- <span>· fov</span> <b id="fov">-</b>
279
- <span>· map</span> <b id="mapspan">-</b></div>
280
- <div class="hud" id="score">
281
- <div class="pts"><b id="total">0</b><span> pts</span></div>
282
- <span id="played">0 episodes</span> · <span id="captured"></span>
283
- </div>
284
- <div class="hud" id="credit"></div>
285
- <div class="hud" id="actions">
286
- <div class="pad" id="padTurn">
287
- <span class="padlabel">look</span>
288
- <div class="btns">
289
- <button id="turn-left" title="Turn 45° left and look (costs 0.01) — key: A or ←">
290
- <span class="glyph">&#9664;</span><span class="tag">left</span></button>
291
- <button id="look-now" title="Look again where you are facing (costs 0.01) — key: L">
292
- <span class="glyph">&#9678;</span><span class="tag">look</span></button>
293
- <button id="turn-right" title="Turn 45° right and look (costs 0.01) — key: D or →">
294
- <span class="glyph">&#9654;</span><span class="tag">right</span></button>
295
- </div>
296
- </div>
297
-
298
- <div class="pad" id="padMove">
299
- <span class="padlabel">walk</span>
300
- <div class="btns">
301
- <button id="move-fwd" title="Walk 15 m forward (costs 0.05) — key: W or ↑">
302
- <span class="glyph">&#9650;</span><span class="tag">forward</span></button>
303
- <button id="move-back" title="Walk 15 m back (costs 0.05) — key: S or ↓">
304
- <span class="glyph">&#9660;</span><span class="tag">back</span></button>
305
- </div>
306
- </div>
307
-
308
- <div class="pad" id="padZoom">
309
- <span class="padlabel">zoom &middot; <b id="fovValue">90&deg;</b></span>
310
- <div class="btns">
311
- <button class="preset" data-fov="90" title="Wide view, 90° (costs 0.01)">wide</button>
312
- <button class="preset" data-fov="50" title="Street level, 50° (costs 0.01)">street</button>
313
- <button class="preset" data-fov="30" title="Read a sign, 30° (costs 0.01)">sign</button>
314
- <input id="fovRange" type="range" min="20" max="110" step="5" value="90"
315
- aria-label="field of view in degrees">
316
- </div>
317
- </div>
318
-
319
- <div class="pad" id="padBudget">
320
- <span class="padlabel">budget</span>
321
- <div class="btns budget">
322
- <b id="stepsBudget">12</b><span>left</span>
323
- <span class="sep">&middot;</span>
324
- <b id="costBudget">0.00</b><span>spent</span>
325
- </div>
326
- </div>
327
- </div>
328
-
329
- <div class="hud" id="expandWrap">
330
- <button id="expand" title="show the rollout trace (T)">+ trace</button>
331
- </div>
332
-
333
- <div id="trace">
334
- <h4>
335
- <span>rollout trace</span>
336
- <span><em id="cost">cost 0.00</em>
337
- <button id="collapse" title="hide the trace (T)">&#8211;</button></span>
338
- </h4>
339
- <div id="steps"></div>
340
- <div id="agentview" style="display:none">
341
- <div class="cap">what the agent sees</div>
342
- <img id="agentimg" alt="the environment's own rendered observation">
343
- </div>
344
- </div>
345
-
346
- <div id="mapwrap">
347
- <div id="map"></div>
348
- <button id="submit">click the map to place a pin</button>
349
- </div>
350
-
351
- <div id="result">
352
- <div class="inner">
353
- <div id="verdict">-</div>
354
- <div class="stat">distance<b id="dist">-</b></div>
355
- <div class="stat">points<b id="points">-</b></div>
356
- <div class="stat env">env reward<b id="reward">-</b></div>
357
- <div class="stat">score - cost<b id="breakdown2">-</b></div>
358
- <div class="stat">true location<b id="truth">-</b></div>
359
- <div id="bar"><i></i></div>
360
- <button id="advance">load another episode</button>
361
- </div>
362
- </div>
363
-
364
- </div>
365
-
366
- <script src="__MAPLIBRE_JS__"></script>
367
- <script src="__PANNELLUM_JS__"></script>
368
- <script>
369
- (function () {
370
- "use strict";
371
- const N_TASKS = __N_TASKS__;
372
- const SPLIT = __SPLIT__;
373
- const MAX_POINTS = __MAX_POINTS__;
374
- // Every task-scoped request has to name its split, or index 12 of eval and
375
- // index 12 of train are indistinguishable and the page reveals the wrong
376
- // ground truth.
377
- const q = (extra) => "?split=" + encodeURIComponent(SPLIT) + (extra || "");
378
- const $ = (id) => document.getElementById(id);
379
-
380
- let socket = null, ready = false, pending = null;
381
- let viewer = null, map = null, guessMarker = null, truthMarker = null;
382
- let lineAdded = false, guess = null, taskIndex = 0, compass = 0;
383
- let taskMeta = null, frameIndex = 0;
384
- let total = 0, played = 0, busy = false, cost = 0;
385
-
386
- const DIRS = ["N", "NE", "E", "SE", "S", "SW", "W", "NW"];
387
- const fmt = (la, lo) => la.toFixed(4) + ", " + lo.toFixed(4);
388
- const yaw = () => (viewer ? ((viewer.getYaw() % 360) + 360) % 360 : 0);
389
- const hfov = () => (viewer ? viewer.getHfov() : 90);
390
-
391
- // ---- the environment, over the same WebSocket session API a client uses --
392
- // Plain REST /step builds a fresh environment per request, so a stateful
393
- // episode has to run over /ws. This page therefore plays exactly the
394
- // rollout an agent would: one reset, a few charged steps, one terminal guess.
395
- function connect() {
396
- const scheme = location.protocol === "https:" ? "wss:" : "ws:";
397
- socket = new WebSocket(scheme + "//" + location.host + "/ws");
398
- socket.onopen = function () {
399
- ready = true;
400
- $("conn").className = "dot live";
401
- $("conn").title = "environment session: connected";
402
- startRound(requestedTask());
403
- };
404
- socket.onclose = function () {
405
- ready = false;
406
- $("conn").className = "dot bad";
407
- $("conn").title = "environment session: disconnected";
408
- };
409
- socket.onerror = function () {
410
- $("conn").className = "dot bad";
411
- $("conn").title = "environment session: error";
412
- };
413
- socket.onmessage = function (event) {
414
- const message = JSON.parse(event.data);
415
- if (message.type === "error") {
416
- addStep("error", message.data ? JSON.stringify(message.data) : "", null);
417
- setBusy(false);
418
- return;
419
- }
420
- if (message.type !== "observation") return;
421
- // The wire format nests the observation and carries reward and done as
422
- // siblings: {observation: {...}, reward, done, metadata}. Flatten it so
423
- // callers read one object.
424
- const payload = message.data || {};
425
- const observation = Object.assign(
426
- {}, payload.observation || payload,
427
- { reward: payload.reward, done: payload.done }
428
- );
429
- const handler = pending;
430
- pending = null;
431
- if (handler) handler(observation);
432
- };
433
- }
434
-
435
- function send(type, data, handler) {
436
- if (!ready) return;
437
- pending = handler || null;
438
- socket.send(JSON.stringify({ type: type, data: data || {} }));
439
- }
440
-
441
- // ---- trace ------------------------------------------------------------
442
- function addStep(op, feedback, observation) {
443
- const row = document.createElement("div");
444
- row.className = "step";
445
- let right = "";
446
- if (observation && observation.steps_remaining !== undefined) {
447
- right = "<span class='rw'>" + observation.steps_remaining + " left</span>";
448
- }
449
- row.innerHTML = "<span class='op'>" + op + "</span>" + right +
450
- "<span class='fb'>" + (feedback || "") + "</span>";
451
- $("steps").appendChild(row);
452
- $("steps").scrollTop = $("steps").scrollHeight;
453
- if (observation && observation.image_base64) {
454
- const mime = observation.image_kind === "map" ? "png" : "jpeg";
455
- $("agentimg").src = "data:image/" + mime + ";base64," + observation.image_base64;
456
- $("agentview").style.display = "block";
457
- }
458
- }
459
-
460
- function applyObservation(observation) {
461
- if (!observation) return;
462
- if (observation.steps_remaining !== undefined) {
463
- $("stepsLeft").textContent = observation.steps_remaining;
464
- $("stepsBudget").textContent = observation.steps_remaining;
465
- }
466
- if (observation.action_cost !== null && observation.action_cost !== undefined) {
467
- cost = observation.action_cost;
468
- $("cost").textContent = "cost " + cost.toFixed(2);
469
- $("costBudget").textContent = cost.toFixed(2);
470
- }
471
- setControls(observation);
472
- }
473
-
474
- // ---- rounds -----------------------------------------------------------
475
- function loadPano(index) {
476
- fetch("/geoguesser/task/" + index + q())
477
- .then((response) => response.json())
478
- .then((meta) => {
479
- taskMeta = meta;
480
- frameIndex = meta.start_frame || 0;
481
- const who = (meta.attribution || {}).creator_username;
482
- $("credit").innerHTML =
483
- "imagery &copy; " + (who ? who : "Mapillary contributor") +
484
- " via <a href='https://www.mapillary.com' target='_blank' rel='noopener'>Mapillary</a>" +
485
- ", <a href='https://creativecommons.org/licenses/by-sa/4.0/' target='_blank' rel='noopener'>CC BY-SA 4.0</a>";
486
- showFrame(frameIndex, 0, 90);
487
- });
488
- }
489
-
490
- /**
491
- * Point the main viewer at one frame of the sequence.
492
- *
493
- * Called on reset and again after every move(), so walking forward actually
494
- * changes what you are looking at rather than only what the trace shows.
495
- * Heading and zoom carry over, because losing your orientation on every step
496
- * would make navigation useless.
497
- */
498
- function showFrame(index, keepYaw, keepHfov) {
499
- frameIndex = index;
500
- const frames = (taskMeta && taskMeta.frames) || [];
501
- const frame = frames[index] || {};
502
- compass = frame.compass_angle || 0;
503
- if (frame.captured_at) {
504
- $("captured").textContent = "captured " + frame.captured_at;
505
- }
506
- $("frame").textContent = index + "/" + Math.max(0, frames.length - 1);
507
- if (viewer) { viewer.destroy(); viewer = null; }
508
- viewer = pannellum.viewer("pano", {
509
- type: "equirectangular",
510
- panorama: "/geoguesser/pano/" + taskIndex + "/" + index + q(),
511
- autoLoad: true, showControls: false, northOffset: compass,
512
- yaw: keepYaw, hfov: keepHfov,
513
- minHfov: 20, maxHfov: 110, compass: false, friction: 0.15,
514
- });
515
- viewer.on("mouseup", updateHud);
516
- viewer.on("touchend", updateHud);
517
- viewer.on("zoomchange", updateHud);
518
- viewer.on("load", updateHud);
519
- setTimeout(updateHud, 400);
520
- }
521
-
522
- function setControls(observation) {
523
- if (!observation) return;
524
- const tools = observation.available_tools || [];
525
- const canLook = tools.indexOf("look") !== -1;
526
- const canMove = tools.indexOf("move") !== -1;
527
- $("padTurn").classList.toggle("gone", !canLook);
528
- $("padZoom").classList.toggle("gone", !canLook);
529
- $("look-now").classList.toggle("gone", !canLook);
530
- $("padMove").classList.toggle(
531
- "gone",
532
- !canMove ||
533
- (!observation.can_move_forward && !observation.can_move_backward)
534
- );
535
- $("move-fwd").classList.toggle("gone", !observation.can_move_forward);
536
- $("move-back").classList.toggle("gone", !observation.can_move_backward);
537
- if (observation.fov_deg) {
538
- const fov = Math.round(observation.fov_deg);
539
- $("fovRange").value = String(fov);
540
- $("fovValue").textContent = fov + "\u00b0";
541
- markPreset(fov);
542
- }
543
- }
544
-
545
- function updateHud() {
546
- if (!viewer) return;
547
- const y = yaw();
548
- $("heading").textContent = DIRS[Math.round(y / 45) % 8] + " " + y.toFixed(0) + "°";
549
- $("fov").textContent = hfov().toFixed(0) + "°";
550
- }
551
-
552
- // ---- charged actions, executed by the environment ---------------------
553
- function setBusy(value) {
554
- busy = value;
555
- $("actions").classList.toggle("working", value);
556
- if (value) {
557
- $("stepsBudget").textContent = "\u2026";
558
- }
559
- }
560
-
561
- function step(op, data, label) {
562
- if (busy || !ready) return;
563
- setBusy(true);
564
- send("step", Object.assign({ op: op }, data), function (observation) {
565
- setBusy(false);
566
- applyObservation(observation);
567
- addStep(label, observation.feedback, observation);
568
- const meta = observation.metadata || {};
569
- if (op === "move" && meta.frame_index !== undefined &&
570
- meta.frame_index !== frameIndex) {
571
- // Keep the player facing the same way through the step.
572
- showFrame(meta.frame_index, yaw(), hfov());
573
- }
574
- if (op === "look" && data && data.fov_deg && viewer) {
575
- // The main view and the agent's view should never disagree.
576
- viewer.setHfov(data.fov_deg);
577
- if (data.heading_deg !== undefined) viewer.setYaw(data.heading_deg);
578
- }
579
- if (op === "guess") reveal(observation);
580
- });
581
- }
582
-
583
- // The pad is the agent's action set, not a viewer control: every button is a
584
- // charged environment step, and the panorama follows the result. Dragging the
585
- // scene stays free, for orientation only.
586
- function lookAt(heading, fov) {
587
- const wrapped = ((Math.round(heading) % 360) + 360) % 360;
588
- step(
589
- "look",
590
- { heading_deg: wrapped, pitch_deg: 0, fov_deg: Math.round(fov) },
591
- "look(heading=" + wrapped + ", fov=" + Math.round(fov) + ")"
592
- );
593
- }
594
-
595
- $("turn-left").onclick = function () { lookAt(yaw() - 45, hfov()); };
596
- $("turn-right").onclick = function () { lookAt(yaw() + 45, hfov()); };
597
- $("look-now").onclick = function () { lookAt(yaw(), hfov()); };
598
-
599
- Array.prototype.forEach.call(
600
- document.querySelectorAll("#padZoom .preset"),
601
- function (button) {
602
- button.onclick = function () {
603
- const fov = parseInt(button.dataset.fov, 10);
604
- $("fovRange").value = String(fov);
605
- $("fovValue").textContent = fov + "\u00b0";
606
- markPreset(fov);
607
- lookAt(yaw(), fov);
608
- };
609
- }
610
- );
611
-
612
- function markPreset(fov) {
613
- Array.prototype.forEach.call(
614
- document.querySelectorAll("#padZoom .preset"),
615
- function (button) {
616
- button.classList.toggle("on", parseInt(button.dataset.fov, 10) === fov);
617
- }
618
- );
619
- }
620
- $("move-fwd").onclick = function () {
621
- step("move", { direction: "forward", meters: 15 }, "move(forward, 15m)");
622
- };
623
- $("move-back").onclick = function () {
624
- step("move", { direction: "backward", meters: 15 }, "move(backward, 15m)");
625
- };
626
-
627
- // The slider reads out live but only spends a step on release, so dragging it
628
- // does not burn the budget.
629
- $("fovRange").addEventListener("input", function () {
630
- const fov = parseInt($("fovRange").value, 10);
631
- $("fovValue").textContent = fov + "\u00b0";
632
- markPreset(fov);
633
- });
634
- $("fovRange").addEventListener("change", function () {
635
- lookAt(yaw(), parseInt($("fovRange").value, 10));
636
- });
637
- function setTrace(visible) {
638
- $("trace").classList.toggle("hidden", !visible);
639
- $("expandWrap").classList.toggle("show", !visible);
640
- }
641
- $("collapse").onclick = function () { setTrace(false); };
642
- $("expand").onclick = function () { setTrace(true); };
643
-
644
- // The picker is the human equivalent of reset(task_index=k): the same call an
645
- // eval harness makes, so a person can replay exactly the episode an agent saw.
646
- /** Task index requested in the page URL, when the host supplied one. */
647
- function requestedTask() {
648
- const value = new URLSearchParams(location.search).get("task");
649
- if (value === null || value === "" || value === "random") return undefined;
650
- const parsed = parseInt(value, 10);
651
- return Number.isFinite(parsed) ? parsed : undefined;
652
- }
653
-
654
- function startRound(index) {
655
- cost = 0;
656
- $("cost").textContent = "cost 0.00";
657
- $("steps").innerHTML = "";
658
- $("agentview").style.display = "none";
659
- clearRound();
660
- const wanted = index === undefined
661
- ? Math.floor(Math.random() * N_TASKS)
662
- : ((index % N_TASKS) + N_TASKS) % N_TASKS;
663
- taskIndex = wanted;
664
- send("reset", { split: SPLIT, index: wanted }, function (observation) {
665
- const meta = observation.metadata || {};
666
- taskIndex = meta.task_index !== undefined ? meta.task_index : wanted;
667
- $("task").textContent = taskIndex;
668
- $("captured").textContent = "captured " + (observation.captured_at || "unknown");
669
- applyObservation(observation);
670
- addStep(
671
- "reset(split='" + SPLIT + "', index=" + taskIndex + ")",
672
- "episode started · " + (observation.available_tools || []).length +
673
- " tools registered",
674
- observation
675
- );
676
- loadPano(taskIndex);
677
- });
678
- }
679
-
680
- // ---- map --------------------------------------------------------------
681
- const WORLD = [[-179, -58], [179, 76]];
682
- map = new maplibregl.Map({
683
- container: "map", style: "__OPENFREEMAP_STYLE__",
684
- center: [0, 12], zoom: 0, minZoom: -2,
685
- attributionControl: { compact: true }, dragRotate: false,
686
- // Without this the world repeats horizontally, which reads as a rendering
687
- // bug at the zoom levels a small guess map uses.
688
- renderWorldCopies: false,
689
- });
690
- map.on("load", () => map.fitBounds(WORLD, { padding: 6, duration: 0 }));
691
- // Exposed so the page can be driven from a test harness or the console.
692
- window.__ggMap = map;
693
- window.__ggState = function () {
694
- return {
695
- busy: busy,
696
- ready: ready,
697
- pendingHandler: !!pending,
698
- socket: socket ? socket.readyState : null,
699
- steps: document.querySelectorAll('#steps .step').length,
700
- };
701
- };
702
-
703
- /**
704
- * Half-width of the visible map, in degrees.
705
- *
706
- * The environment renders its own map from this, so a pin dropped while
707
- * zoomed into a city comes back as a street-level map rather than a
708
- * continental one. Without it the agent's view and the player's would
709
- * disagree about how precisely the pin could be aimed.
710
- */
711
- function currentSpanDeg() {
712
- const bounds = map.getBounds();
713
- const span = Math.abs(bounds.getEast() - bounds.getWest()) / 2;
714
- return Math.min(180, Math.max(0.03, span));
715
- }
716
-
717
- function updateMapSpan() {
718
- const span = currentSpanDeg();
719
- $("mapspan").textContent =
720
- span >= 1 ? span.toFixed(0) + "\u00b0" : (span * 111).toFixed(0) + " km";
721
- }
722
- map.on("zoomend", updateMapSpan);
723
- map.on("moveend", updateMapSpan);
724
- map.on("load", updateMapSpan);
725
-
726
- map.on("click", function (event) {
727
- if (busy || $("result").classList.contains("show")) return;
728
- // Once you have committed to a pin the map stays open; letting it collapse
729
- // on mouse-out makes it easy to lose the guess you were adjusting.
730
- $("mapwrap").classList.add("big");
731
- map.resize();
732
- guess = event.lngLat;
733
- if (guessMarker) guessMarker.remove();
734
- guessMarker = new maplibregl.Marker({ color: "#c4332a" })
735
- .setLngLat(guess).addTo(map);
736
- $("submit").className = "ready";
737
- $("submit").textContent = "submit guess";
738
- // A pin is a real, charged environment step, so the map the agent would see
739
- // comes back in the trace panel — framed at the zoom you are looking at, so
740
- // the two views agree about how precisely the pin was aimed.
741
- const span = currentSpanDeg();
742
- step("pin", { lat: guess.lat, lon: guess.lng, span_deg: span },
743
- "place_pin(" + guess.lat.toFixed(2) + ", " + guess.lng.toFixed(2) +
744
- ", span=" + span.toFixed(2) + ")");
745
- });
746
-
747
- $("submit").onclick = function () {
748
- if (!guess) return;
749
- step("guess", { lat: guess.lat, lon: guess.lng },
750
- "submit_guess(" + guess.lat.toFixed(2) + ", " + guess.lng.toFixed(2) + ")");
751
- };
752
-
753
- function reveal(observation) {
754
- if (observation.distance_km === null || observation.distance_km === undefined) {
755
- addStep("guess rejected",
756
- observation.feedback || "the environment returned no distance",
757
- observation);
758
- return;
759
- }
760
- const km = observation.distance_km;
761
- const reward = observation.reward === null ? 0 : observation.reward;
762
- const score = observation.score === null ? 0 : observation.score;
763
- const points = Math.round(score * MAX_POINTS);
764
- const truthLat = observation.true_lat, truthLon = observation.true_lon;
765
- total += points;
766
- played += 1;
767
- $("played").textContent = played + (played === 1 ? " episode" : " episodes");
768
-
769
- $("verdict").textContent =
770
- km < 0.025 ? "Perfect." : km < 25 ? "Pinpoint." : km < 200 ? "Close." :
771
- km < 1500 ? "Right region." : "Wrong continent.";
772
- $("dist").textContent = km < 10 ? (km * 1000).toFixed(0) + " m" : km.toFixed(0) + " km";
773
- $("points").textContent = points + " / " + MAX_POINTS;
774
- $("reward").textContent = reward.toFixed(3);
775
- $("breakdown2").textContent =
776
- score.toFixed(3) + " - " + (observation.action_cost || 0).toFixed(2);
777
- $("truth").textContent = fmt(truthLat, truthLon);
778
- $("bar").firstElementChild.style.width = (reward * 100).toFixed(1) + "%";
779
- $("total").textContent = total;
780
- $("result").classList.add("show");
781
- document.body.classList.add("revealing");
782
- $("actions").style.opacity = "0";
783
-
784
- truthMarker = new maplibregl.Marker({ color: "#5aa06e" })
785
- .setLngLat([truthLon, truthLat]).addTo(map);
786
- const line = {
787
- type: "Feature",
788
- geometry: {
789
- type: "LineString",
790
- coordinates: [[guess.lng, guess.lat], [truthLon, truthLat]],
791
- },
792
- };
793
- if (lineAdded) {
794
- map.getSource("shot").setData(line);
795
- } else {
796
- map.addSource("shot", { type: "geojson", data: line });
797
- map.addLayer({
798
- id: "shot", type: "line", source: "shot",
799
- paint: { "line-color": "#c4332a", "line-width": 2.5, "line-dasharray": [2, 1.6] },
800
- });
801
- lineAdded = true;
802
- }
803
- $("mapwrap").classList.add("reveal");
804
- map.resize();
805
- setTimeout(function () {
806
- map.fitBounds(
807
- [[Math.min(guess.lng, truthLon), Math.min(guess.lat, truthLat)],
808
- [Math.max(guess.lng, truthLon), Math.max(guess.lat, truthLat)]],
809
- { padding: 80, maxZoom: 7, duration: 900 }
810
- );
811
- }, 260);
812
- }
813
-
814
- function clearRound() {
815
- guess = null;
816
- setBusy(false);
817
- if (guessMarker) { guessMarker.remove(); guessMarker = null; }
818
- if (truthMarker) { truthMarker.remove(); truthMarker = null; }
819
- if (lineAdded) {
820
- map.getSource("shot").setData({
821
- type: "Feature", geometry: { type: "LineString", coordinates: [] },
822
- });
823
- }
824
- $("mapwrap").classList.remove("reveal", "big");
825
- map.resize();
826
- map.fitBounds(WORLD, { padding: 6, duration: 700 });
827
- $("submit").className = "";
828
- $("submit").textContent = "click the map to place a pin";
829
- $("result").classList.remove("show");
830
- document.body.classList.remove("revealing");
831
- $("actions").style.opacity = "1";
832
- }
833
-
834
- // An episode is one guess, so there is no round to advance: the terminal
835
- // control simply starts another episode.
836
- $("advance").onclick = function () {
837
- clearRound();
838
- startRound();
839
- };
840
-
841
- document.addEventListener("keydown", function (event) {
842
- if (event.key === "m" || event.key === "M") {
843
- $("mapwrap").classList.toggle("big");
844
- map.resize();
845
- if (!guess && !busy) map.fitBounds(WORLD, { padding: 6, duration: 250 });
846
- } else if (event.key === "t" || event.key === "T") {
847
- setTrace($("trace").classList.contains("hidden"));
848
- } else if (event.key === "ArrowUp" || event.key === "w") {
849
- if (!$("move-fwd").classList.contains("gone")) $("move-fwd").click();
850
- } else if (event.key === "ArrowDown" || event.key === "s") {
851
- if (!$("move-back").classList.contains("gone")) $("move-back").click();
852
- } else if (event.key === "l" || event.key === "L") {
853
- if (!$("padTurn").classList.contains("gone")) $("look-now").click();
854
- } else if (event.key === "ArrowLeft" || event.key === "a") {
855
- if (!$("padTurn").classList.contains("gone")) $("turn-left").click();
856
- } else if (event.key === "ArrowRight" || event.key === "d") {
857
- if (!$("padTurn").classList.contains("gone")) $("turn-right").click();
858
- } else if (event.key === "Enter") {
859
- if ($("result").classList.contains("show")) { $("advance").click(); }
860
- else { $("submit").click(); }
861
- }
862
- });
863
-
864
- connect();
865
- })();
866
- </script>
867
- </body>
868
- </html>
869
- """
870
-
871
-
872
- def play_page_html(splits: list[dict] | int, split: str | None = None) -> str:
873
- """
874
- Return the standalone play page.
875
-
876
- Served at `/geoguesser/play` and embedded in the Gradio tab through an
877
- iframe. It is a full document rather than a fragment because `gr.HTML`
878
- inserts markup without running `<script>` tags.
879
-
880
- Args:
881
- splits (`list[dict]` or `int`):
882
- Split descriptors from [`~GeoGuesserEnvironment.list_splits`]. A
883
- bare integer is accepted as a task count for callers predating
884
- splits.
885
- split (`str`, *optional*):
886
- Which split the page should play. Defaults to the split marked
887
- `default`, else the first one.
888
-
889
- Returns:
890
- `str`: A complete HTML document.
891
- """
892
- if isinstance(splits, int):
893
- descriptors = [
894
- {"name": "train", "num_tasks": splits, "default": True, "type": "train"}
895
- ]
896
- else:
897
- descriptors = list(splits) or [
898
- {"name": "train", "num_tasks": 1, "default": True, "type": "train"}
899
- ]
900
- chosen = next(
901
- (d for d in descriptors if d["name"] == split),
902
- next((d for d in descriptors if d.get("default")), descriptors[0]),
903
- )
904
- replacements = {
905
- "__SPLIT__": json.dumps(chosen["name"]),
906
- "__N_TASKS__": str(max(1, int(chosen.get("num_tasks", 1)))),
907
- "__MAX_POINTS__": str(MAX_POINTS_PER_ROUND),
908
- "__MAPLIBRE_JS__": MAPLIBRE_JS,
909
- "__MAPLIBRE_CSS__": MAPLIBRE_CSS,
910
- "__PANNELLUM_JS__": PANNELLUM_JS,
911
- "__PANNELLUM_CSS__": PANNELLUM_CSS,
912
- "__OPENFREEMAP_STYLE__": OPENFREEMAP_STYLE,
913
- }
914
- page = _TEMPLATE
915
- for token, value in replacements.items():
916
- page = page.replace(token, value)
917
- return page
918
-
919
-
920
- def _iframe(task: str | int = "random", split: str = "") -> str:
921
- """Markup for the play iframe, pointed at one task of one split.
922
-
923
- Args:
924
- task (`str` or `int`, *optional*, defaults to `"random"`):
925
- Task index to open, or `"random"`.
926
- split (`str`, *optional*):
927
- Split to play. Empty means the server's default split.
928
-
929
- Returns:
930
- `str`: An iframe element. Re-rendering it with a different task is what
931
- makes the Gradio controls reload the round, since the page reads its
932
- task from the URL.
933
- """
934
- query = f"?task={task}"
935
- if split:
936
- query += f"&split={urllib.parse.quote(split)}"
937
- return (
938
- f'<iframe src="/geoguesser/play{query}" '
939
- 'style="width:100%;height:720px;border:1px solid #2c353d;'
940
- 'border-radius:6px" allow="fullscreen"></iframe>'
941
- )
942
-
943
-
944
- def build_geoguesser_gradio_app(
945
- web_manager: Any,
946
- action_fields: List[Dict[str, Any]],
947
- metadata: Optional[Any],
948
- is_chat_env: bool,
949
- title: str,
950
- quick_start_md: str,
951
- ) -> gr.Blocks:
952
- """
953
- Build the human-play tab.
954
-
955
- The episode controls live here, on the Gradio side, rather than inside the
956
- page: picking a task is orchestration, the same `reset(task_index=k)` an
957
- eval harness calls, so it belongs with the host controls and not among the
958
- in-game HUD.
959
-
960
- Args:
961
- web_manager (`Any`):
962
- The playground's environment manager, unused here.
963
- action_fields (`list[dict]`):
964
- Action schema fields, unused here.
965
- metadata (`Any`, *optional*):
966
- Environment metadata, unused here.
967
- is_chat_env (`bool`):
968
- Whether the env is chat-shaped, unused here.
969
- title (`str`):
970
- Playground title.
971
- quick_start_md (`str`):
972
- Quick-start markdown, unused here.
973
-
974
- Returns:
975
- `gradio.Blocks`: The play tab, hosting `/geoguesser/play` in an iframe.
976
- """
977
- # Ask the server which splits it actually serves, rather than re-deriving
978
- # them here from environment variables and drifting out of step with it.
979
- descriptors: list[dict] = []
980
- try:
981
- from .app import ACTIVE_DEFAULT_SPLIT, create_geoguesser_environment
982
-
983
- descriptors = create_geoguesser_environment().list_splits()
984
- default_split = ACTIVE_DEFAULT_SPLIT
985
- except Exception: # pragma: no cover - the page still works without counts
986
- default_split = "train"
987
- if not descriptors:
988
- descriptors = [
989
- {"name": default_split, "num_tasks": 1, "default": True, "type": "train"}
990
- ]
991
-
992
- counts = {d["name"]: max(1, int(d.get("num_tasks", 1))) for d in descriptors}
993
- names = list(counts)
994
- if default_split not in counts:
995
- default_split = names[0]
996
-
997
- def _label(split: str) -> str:
998
- return f"reset(index=) · 0 to {counts[split] - 1}"
999
-
1000
- with gr.Blocks(title="Geoguesser Environment") as blocks:
1001
- with gr.Row():
1002
- split_box = gr.Dropdown(
1003
- choices=names,
1004
- value=default_split,
1005
- label="reset(split=)",
1006
- scale=1,
1007
- interactive=len(names) > 1,
1008
- )
1009
- task_box = gr.Number(
1010
- value=0,
1011
- minimum=0,
1012
- maximum=counts[default_split] - 1,
1013
- step=1,
1014
- precision=0,
1015
- label=_label(default_split),
1016
- scale=2,
1017
- )
1018
- load_button = gr.Button("load episode", variant="primary", scale=1)
1019
- random_button = gr.Button("random episode", scale=1)
1020
- frame = gr.HTML(value=_iframe("random", default_split), show_label=False)
1021
-
1022
- def _on_split(split: str):
1023
- """Re-range the index box so it cannot address a missing task."""
1024
- split = split or default_split
1025
- return gr.update(maximum=counts[split] - 1, value=0, label=_label(split))
1026
-
1027
- split_box.change(fn=_on_split, inputs=split_box, outputs=task_box)
1028
- load_button.click(
1029
- fn=lambda index, split: _iframe(int(index or 0), split or default_split),
1030
- inputs=[task_box, split_box],
1031
- outputs=frame,
1032
- )
1033
- random_button.click(
1034
- fn=lambda split: _iframe(
1035
- random.randrange(counts[split or default_split]),
1036
- split or default_split,
1037
- ),
1038
- inputs=split_box,
1039
- outputs=frame,
1040
- )
1041
- return blocks
1042
-
1043
-
1044
- __all__ = ["build_geoguesser_gradio_app", "play_page_html"]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/parser.py DELETED
@@ -1,149 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Extract coordinates from a model's free-text reply.
4
-
5
- Models emit reasoning and coordinates together, in many shapes. The parser
6
- accepts what they actually produce rather than demanding a schema, and returns
7
- `None` when nothing usable is present so the failure lands in the reward
8
- instead of raising.
9
- """
10
-
11
- from __future__ import annotations
12
-
13
- import re
14
- from dataclasses import dataclass
15
-
16
-
17
- # 48.8584, 2.2945 | -16.49 / -68.12 | lat: 12.9 lon: 77.5
18
- _DECIMAL_PAIR = re.compile(
19
- r"(-?\d{1,3}(?:\.\d+)?)\s*(?:,|/|;|\s+and\s+|\s+)\s*(-?\d{1,3}(?:\.\d+)?)"
20
- )
21
-
22
- # 48°51'29"N 2°17'40"E
23
- _DMS = re.compile(
24
- r"(\d{1,3})\s*[°d]\s*(\d{1,2})?\s*['′m]?\s*(\d{1,2}(?:\.\d+)?)?"
25
- r"\s*[\"″s]?\s*([NSEW])",
26
- re.IGNORECASE,
27
- )
28
-
29
- _LABELLED = re.compile(
30
- r"lat(?:itude)?\s*[:=]\s*(-?\d{1,3}(?:\.\d+)?)"
31
- r".{0,40}?"
32
- r"lon(?:g|gitude)?\s*[:=]\s*(-?\d{1,3}(?:\.\d+)?)",
33
- re.IGNORECASE | re.DOTALL,
34
- )
35
-
36
- _TAG = re.compile(r"<guess>(.*?)</guess>", re.IGNORECASE | re.DOTALL)
37
-
38
- _JSON_ISH = re.compile(
39
- r"\"lat(?:itude)?\"\s*:\s*(-?\d{1,3}(?:\.\d+)?)"
40
- r".{0,60}?"
41
- r"\"lon(?:g|gitude)?\"\s*:\s*(-?\d{1,3}(?:\.\d+)?)",
42
- re.IGNORECASE | re.DOTALL,
43
- )
44
-
45
-
46
- @dataclass
47
- class ParsedGuess:
48
- """Outcome of parsing a reply.
49
-
50
- Attributes:
51
- lat (`float` or `None`):
52
- Latitude, or `None` when nothing could be extracted.
53
- lon (`float` or `None`):
54
- Longitude, or `None` when nothing could be extracted.
55
- source (`str`):
56
- Which pattern matched: `"tag"`, `"json"`, `"labelled"`, `"dms"`,
57
- `"decimal"` or `"none"`.
58
- note (`str`):
59
- Short explanation, safe to show the model as feedback.
60
- """
61
-
62
- lat: float | None
63
- lon: float | None
64
- source: str
65
- note: str = ""
66
-
67
- @property
68
- def ok(self) -> bool:
69
- """Whether a usable coordinate pair was extracted."""
70
- return self.lat is not None and self.lon is not None
71
-
72
-
73
- def _valid(lat: float, lon: float) -> bool:
74
- return -90.0 <= lat <= 90.0 and -180.0 <= lon <= 180.0
75
-
76
-
77
- def _dms_to_decimal(deg: str, minute: str | None, sec: str | None, hemi: str) -> float:
78
- value = float(deg) + float(minute or 0) / 60 + float(sec or 0) / 3600
79
- return -value if hemi.upper() in ("S", "W") else value
80
-
81
-
82
- def parse_guess(response: str) -> ParsedGuess:
83
- """
84
- Pull a coordinate pair out of a model reply.
85
-
86
- Patterns are tried most explicit first, so a `<guess>` tag or a labelled
87
- `lat:`/`lon:` pair wins over a bare number pair that might be a date or a
88
- step count.
89
-
90
- Args:
91
- response (`str`):
92
- The model's unedited reply.
93
-
94
- Returns:
95
- [`ParsedGuess`]: The extracted coordinates, or a result whose `ok` is
96
- `False` with a `note` explaining what was wrong.
97
-
98
- Examples:
99
-
100
- ```python
101
- parse_guess("I think coastal Portugal. <guess>38.72, -9.14</guess>")
102
- ```
103
- """
104
- if not response or not response.strip():
105
- return ParsedGuess(None, None, "none", "Empty response.")
106
-
107
- tagged = _TAG.search(response)
108
- haystacks = [(tagged.group(1), "tag")] if tagged else []
109
- haystacks.append((response, "body"))
110
-
111
- for text, origin in haystacks:
112
- for pattern, name in ((_JSON_ISH, "json"), (_LABELLED, "labelled")):
113
- m = pattern.search(text)
114
- if m:
115
- lat, lon = float(m.group(1)), float(m.group(2))
116
- if _valid(lat, lon):
117
- src = name if origin == "body" else "tag"
118
- return ParsedGuess(lat, lon, src)
119
- return ParsedGuess(
120
- None, None, "none", f"Coordinates out of range: {lat}, {lon}."
121
- )
122
-
123
- dms = _DMS.findall(text)
124
- if len(dms) >= 2:
125
- lat_m = next((d for d in dms if d[3].upper() in ("N", "S")), None)
126
- lon_m = next((d for d in dms if d[3].upper() in ("E", "W")), None)
127
- if lat_m and lon_m:
128
- lat = _dms_to_decimal(*lat_m)
129
- lon = _dms_to_decimal(*lon_m)
130
- if _valid(lat, lon):
131
- return ParsedGuess(lat, lon, "dms")
132
-
133
- m = _DECIMAL_PAIR.search(text)
134
- if m:
135
- lat, lon = float(m.group(1)), float(m.group(2))
136
- if _valid(lat, lon):
137
- src = "decimal" if origin == "body" else "tag"
138
- return ParsedGuess(lat, lon, src)
139
- return ParsedGuess(
140
- None, None, "none", f"Coordinates out of range: {lat}, {lon}."
141
- )
142
-
143
- return ParsedGuess(
144
- None,
145
- None,
146
- "none",
147
- "No coordinates found. Reply with a latitude and longitude, for "
148
- "example <guess>48.8584, 2.2945</guess>.",
149
- )
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/render/__init__.py DELETED
@@ -1,3 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Image rendering for the GeoGuesser environment."""
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/render/minimap.py DELETED
@@ -1,824 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """The guess map, rendered offline.
4
-
5
- Everything here runs against bundled Natural Earth vectors, so a pin costs no
6
- network call and always renders the same bytes. Map detail is a function of
7
- zoom alone and never of proximity to the target: prefetching finer data around
8
- task locations would turn the map into a ground-truth oracle.
9
-
10
- The renderer answers only "where did the agent point?". It never draws, names
11
- or hints at the true location.
12
- """
13
-
14
- from __future__ import annotations
15
-
16
- import functools
17
- import hashlib
18
- import io
19
- import json
20
- import logging
21
- import math
22
- import os
23
- import pathlib
24
- import urllib.parse
25
- import urllib.request
26
- from dataclasses import dataclass
27
-
28
- import matplotlib
29
-
30
- matplotlib.use("Agg")
31
-
32
- import matplotlib.patheffects as path_effects # noqa: E402
33
- import matplotlib.pyplot as plt # noqa: E402
34
- from matplotlib.collections import LineCollection # noqa: E402
35
- from matplotlib.patches import Polygon as MplPolygon # noqa: E402
36
- from matplotlib.path import Path as MplPath # noqa: E402
37
- from PIL import Image # noqa: E402
38
-
39
- from ..scoring import haversine_km as _haversine_km # noqa: E402
40
-
41
- logger = logging.getLogger(__name__)
42
-
43
-
44
- GEO_DIR = pathlib.Path(__file__).resolve().parents[2] / "data" / "geo"
45
- COUNTRIES_FILE = GEO_DIR / "ne_110m_admin_0_countries.geojson"
46
- CITIES_FILE = GEO_DIR / "ne_50m_populated_places.geojson"
47
-
48
- # Optional detail layers, fetched by scripts/fetch_detail_geo.py. Without them
49
- # the map shows outlines and major cities only, which is enough to place a
50
- # country but not to aim within a city — while the human's map is street-level.
51
- # The reward measures precision, so the two views have to agree about how
52
- # precisely a pin can be aimed.
53
- DETAIL_DIR = GEO_DIR / "detail"
54
-
55
- # Detail appears as a function of zoom alone, never of proximity to the answer.
56
- # Loading finer data near task locations would turn the cache into an oracle.
57
- # Below this the map fetches real OSM ways, because Natural Earth 10m tops out
58
- # at highway level: it will show the motorways around a city but not the street
59
- # grid inside it, and the player's tiles show both. Responses are cached to
60
- # disk, so the first render of a neighbourhood is slow and every later one is
61
- # instant and byte-identical.
62
- STREET_MAX_SPAN_DEG = 0.35
63
- STREET_LABEL_MAX_SPAN_DEG = 0.08
64
- """Street names only below this span. Wider, and the names collide."""
65
- MAX_STREET_LABELS = 14
66
- """A cap, because a dense grid has hundreds of named ways."""
67
- OVERPASS_URL = "https://overpass-api.de/api/interpreter"
68
- OVERPASS_TIMEOUT_S = 40.0
69
- OSM_CACHE_DIR = pathlib.Path(
70
- os.environ.get("GEOGUESSER_OSM_CACHE", str(GEO_DIR / "osm_cache"))
71
- )
72
- """Where fetched street windows are cached.
73
-
74
- Configurable so a deployment that cannot reach Overpass -- a Space, whose
75
- datacenter egress gets 504s from the public instance -- can mount a pre-warmed
76
- cache instead of silently rendering maps without streets.
77
- """
78
-
79
- OVERPASS_ATTEMPTS = 2
80
- """Overpass 504s under load; one cheap retry recovers a useful fraction."""
81
-
82
- _STREET_FETCH_FAILED = [False]
83
- """Whether the most recent street fetch failed, for observation metadata.
84
-
85
- A map quietly missing its streets looks like a styling choice rather than a
86
- degraded environment, so the failure is reported rather than inferred.
87
- """
88
-
89
-
90
- def street_fetch_failed() -> bool:
91
- """Whether the last street fetch in this process failed."""
92
- return _STREET_FETCH_FAILED[0]
93
-
94
-
95
- _OSM_SCHEMA = 2
96
- """Bumped whenever the per-way record changes, to retire stale caches."""
97
-
98
- # Ways worth drawing, heaviest first, with the line width to draw them at.
99
- STREET_WEIGHTS = {
100
- "motorway": 2.2,
101
- "trunk": 2.0,
102
- "primary": 1.7,
103
- "secondary": 1.4,
104
- "tertiary": 1.1,
105
- "residential": 0.8,
106
- "unclassified": 0.8,
107
- "living_street": 0.7,
108
- "service": 0.5,
109
- "pedestrian": 0.5,
110
- "motorway_link": 1.2,
111
- "trunk_link": 1.1,
112
- "primary_link": 1.0,
113
- "secondary_link": 0.9,
114
- "tertiary_link": 0.8,
115
- }
116
-
117
- URBAN_MAX_SPAN_DEG = 4.0
118
- ROADS_MAX_SPAN_DEG = 4.0
119
- RIVERS_MAX_SPAN_DEG = 6.0
120
- TOWNS_MAX_SPAN_DEG = 4.0
121
-
122
- # Version of the bundled geodata. Any change alters reverse-geocode text and
123
- # therefore the observations, so eval scores must cite it.
124
- GEODATA_VERSION = "natural-earth-110m+50m/2024.1"
125
-
126
- _LAND = "#e9e5dd"
127
- _ROAD = "#8a6a52"
128
- _RIVER = "#6f9fba"
129
- _URBAN = "#e4ded6"
130
- _TOWN = "#3d3833"
131
- _WATER = "#cfe0ea"
132
- _BORDER = "#5f5a54"
133
- _PIN = "#c4332a"
134
- _TRUTH = "#3f7a55"
135
- _INK = "#2b2a28"
136
-
137
- _COMPASS = ("N", "NE", "E", "SE", "S", "SW", "W", "NW")
138
-
139
- # Mutable so the server can turn street fetching off for a run that must never
140
- # touch the network, without threading a flag through every render call.
141
- _STREETS_ENABLED = [True]
142
-
143
-
144
- def set_street_detail(enabled: bool) -> None:
145
- """Turn OSM street fetching on or off for this process."""
146
- _STREETS_ENABLED[0] = bool(enabled)
147
-
148
-
149
- def street_detail_enabled() -> bool:
150
- """Whether OSM street fetching is currently on."""
151
- return _STREETS_ENABLED[0]
152
-
153
-
154
- @dataclass(frozen=True)
155
- class Place:
156
- """Where a coordinate falls, in words.
157
-
158
- Attributes:
159
- country (`str` or `None`):
160
- Country name, or `None` over open water.
161
- continent (`str` or `None`):
162
- Continent name, or `None` over open water.
163
- subregion (`str` or `None`):
164
- UN subregion, or `None` over open water.
165
- nearest_city (`str`):
166
- Name of the closest populated place in the bundled dataset.
167
- city_distance_km (`float`):
168
- Distance to that city in kilometres.
169
- city_bearing (`str`):
170
- Compass direction from the city to the coordinate.
171
- """
172
-
173
- country: str | None
174
- continent: str | None
175
- subregion: str | None
176
- nearest_city: str
177
- city_distance_km: float
178
- city_bearing: str
179
-
180
-
181
- @functools.lru_cache(maxsize=1)
182
- def _countries() -> list[dict]:
183
- return json.loads(COUNTRIES_FILE.read_text())["features"]
184
-
185
-
186
- @functools.lru_cache(maxsize=1)
187
- def _cities() -> list[tuple[str, float, float]]:
188
- feats = json.loads(CITIES_FILE.read_text())["features"]
189
- out = []
190
- for f in feats:
191
- lon, lat = f["geometry"]["coordinates"]
192
- out.append((f["properties"]["name"], lat, lon))
193
- return out
194
-
195
-
196
- @functools.lru_cache(maxsize=8)
197
- def _detail(layer: str) -> list[dict]:
198
- """Load one optional detail layer, or an empty list when absent."""
199
- path = DETAIL_DIR / f"{layer}.json"
200
- if not path.exists():
201
- return []
202
- return json.loads(path.read_text())
203
-
204
-
205
- def _osm_cache_path(lat: float, lon: float, span: float) -> pathlib.Path:
206
- """Cache file for one street window, keyed by a quantised bounding box.
207
-
208
- The key carries `_OSM_SCHEMA`, so widening what is stored per way retires
209
- the old files instead of serving geometry with no labels.
210
- """
211
- key = f"v{_OSM_SCHEMA}_{round(lat, 3):.3f}_{round(lon, 3):.3f}_{round(span, 4):.4f}"
212
- digest = hashlib.sha1(key.encode()).hexdigest()[:16]
213
- return OSM_CACHE_DIR / f"{digest}.json"
214
-
215
-
216
- def _fetch_streets(lat: float, lon: float, span: float) -> list[dict]:
217
- """Ask Overpass for the ways in one window, or return an empty list.
218
-
219
- Any failure — offline, rate limited, malformed — degrades to no streets
220
- rather than failing the step that asked for the map.
221
- """
222
- box = f"{lat - span},{lon - span},{lat + span},{lon + span}"
223
- # Street names and road numbers are what make the agent's map comparable to
224
- # the human's, and they cost nothing extra: Overpass returns tags with the
225
- # geometry either way.
226
- query = f'[out:json][timeout:30];way["highway"]({box});out tags geom;'
227
- payload = None
228
- for attempt in range(1, OVERPASS_ATTEMPTS + 1):
229
- request = urllib.request.Request(
230
- OVERPASS_URL,
231
- data=urllib.parse.urlencode({"data": query}).encode(),
232
- headers={"User-Agent": "openenv-geoguesser-env/0.1 (research environment)"},
233
- )
234
- try:
235
- with urllib.request.urlopen(
236
- request, timeout=OVERPASS_TIMEOUT_S
237
- ) as response:
238
- payload = json.loads(response.read())
239
- break
240
- except Exception as exc: # noqa: BLE001 - a missing map must not end a step
241
- logger.warning(
242
- "street detail unavailable for %.3f,%.3f (attempt %d/%d): %r",
243
- lat,
244
- lon,
245
- attempt,
246
- OVERPASS_ATTEMPTS,
247
- exc,
248
- )
249
- if payload is None:
250
- _STREET_FETCH_FAILED[0] = True
251
- return []
252
- _STREET_FETCH_FAILED[0] = False
253
- ways = []
254
- for element in payload.get("elements", []):
255
- geometry = element.get("geometry") or []
256
- if len(geometry) < 2:
257
- continue
258
- tags = element.get("tags") or {}
259
- kind = tags.get("highway", "residential")
260
- way = {
261
- "w": STREET_WEIGHTS.get(kind, 0.6),
262
- "c": [[point["lon"], point["lat"]] for point in geometry],
263
- }
264
- # A road number is often the only label a rural road has, and it is
265
- # exactly the clue a player reads off a sign.
266
- label = tags.get("name") or tags.get("ref")
267
- if label:
268
- way["n"] = label[:34]
269
- ways.append(way)
270
- return ways
271
-
272
-
273
- def street_ways(lat: float, lon: float, span: float) -> list[dict]:
274
- """
275
- Street geometry for one window, cached on disk.
276
-
277
- Args:
278
- lat (`float`):
279
- Latitude at the centre of the window.
280
- lon (`float`):
281
- Longitude at the centre of the window.
282
- span (`float`):
283
- Half-width of the window in degrees.
284
-
285
- Returns:
286
- `list[dict]`: One entry per way, with a line width `w` and a coordinate
287
- list `c`. Empty when the window is too wide, streets are disabled, or
288
- the fetch failed.
289
- """
290
- path = _osm_cache_path(lat, lon, span)
291
- if path.exists():
292
- try:
293
- return json.loads(path.read_text())
294
- except json.JSONDecodeError:
295
- path.unlink(missing_ok=True)
296
- ways = _fetch_streets(lat, lon, span)
297
- OSM_CACHE_DIR.mkdir(parents=True, exist_ok=True)
298
- path.write_text(json.dumps(ways, separators=(",", ":")))
299
- logger.info(
300
- "cached %d street ways for %.3f,%.3f span %.3f", len(ways), lat, lon, span
301
- )
302
- return ways
303
-
304
-
305
- def has_detail() -> bool:
306
- """Whether the optional detail layers are installed."""
307
- return any(
308
- (DETAIL_DIR / f"{name}.json").exists()
309
- for name in ("roads", "urban", "places", "rivers")
310
- )
311
-
312
-
313
- def _parts(row: dict) -> list[list]:
314
- """Coordinate lists for one compacted feature, whatever its geometry."""
315
- kind, coordinates = row["t"], row["c"]
316
- if kind in ("LineString", "Point"):
317
- return [coordinates] if kind == "LineString" else [[coordinates]]
318
- if kind == "MultiLineString":
319
- return coordinates
320
- if kind == "Polygon":
321
- return [coordinates[0]]
322
- if kind == "MultiPolygon":
323
- return [polygon[0] for polygon in coordinates]
324
- return []
325
-
326
-
327
- def _near(points: list, lat: float, lon: float, span: float) -> bool:
328
- return any(
329
- abs(point[0] - lon) < span and abs(point[1] - lat) < span for point in points
330
- )
331
-
332
-
333
- def _rings(geometry: dict) -> list[list]:
334
- if geometry["type"] == "Polygon":
335
- return [geometry["coordinates"][0]]
336
- if geometry["type"] == "MultiPolygon":
337
- return [poly[0] for poly in geometry["coordinates"]]
338
- return []
339
-
340
-
341
- def _bearing(from_lat: float, from_lon: float, to_lat: float, to_lon: float) -> str:
342
- d_lon = math.radians(to_lon - from_lon)
343
- lat_a, lat_b = math.radians(from_lat), math.radians(to_lat)
344
- y = math.sin(d_lon) * math.cos(lat_b)
345
- x = math.cos(lat_a) * math.sin(lat_b) - math.sin(lat_a) * math.cos(
346
- lat_b
347
- ) * math.cos(d_lon)
348
- deg = (math.degrees(math.atan2(y, x)) + 360) % 360
349
- return _COMPASS[int(deg / 45 + 0.5) % 8]
350
-
351
-
352
- def locate(lat: float, lon: float) -> Place:
353
- """
354
- Describe a coordinate using only the bundled vectors.
355
-
356
- Args:
357
- lat (`float`):
358
- Latitude in degrees.
359
- lon (`float`):
360
- Longitude in degrees.
361
-
362
- Returns:
363
- [`Place`]: Administrative names for the point, plus the nearest
364
- populated place with distance and bearing.
365
- """
366
- country = continent = subregion = None
367
- for feature in _countries():
368
- for ring in _rings(feature["geometry"]):
369
- if MplPath(ring).contains_point((lon, lat)):
370
- props = feature["properties"]
371
- country = props.get("ADMIN")
372
- continent = props.get("CONTINENT")
373
- subregion = props.get("SUBREGION")
374
- break
375
- if country:
376
- break
377
-
378
- best_name, best_km, best_bearing = "", float("inf"), "N"
379
- for name, city_lat, city_lon in _cities():
380
- km = _haversine_km(lat, lon, city_lat, city_lon)
381
- if km < best_km:
382
- best_name, best_km = name, km
383
- best_bearing = _bearing(city_lat, city_lon, lat, lon)
384
- return Place(country, continent, subregion, best_name, best_km, best_bearing)
385
-
386
-
387
- def describe_pin(
388
- index: int, lat: float, lon: float, previous: tuple[float, float] | None = None
389
- ) -> str:
390
- """
391
- One line of feedback for a pin, containing nothing about the target.
392
-
393
- Args:
394
- index (`int`):
395
- 1-based pin number.
396
- lat (`float`):
397
- Latitude of the pin.
398
- lon (`float`):
399
- Longitude of the pin.
400
- previous (`tuple[float, float]`, *optional*):
401
- The preceding pin, so the agent can compare its own candidates.
402
-
403
- Returns:
404
- `str`: Feedback describing the pinned location.
405
- """
406
- place = locate(lat, lon)
407
- where = (
408
- f"{place.country} ({place.subregion})"
409
- if place.country
410
- else "open water - no landmass at this coordinate"
411
- )
412
- line = (
413
- f"Pin {index} placed at {lat:.4f}, {lon:.4f} - {where}. "
414
- f"Nearest major city: {place.nearest_city}, "
415
- f"~{place.city_distance_km:.0f} km {place.city_bearing}."
416
- )
417
- if previous is not None:
418
- km = _haversine_km(lat, lon, previous[0], previous[1])
419
- line += f" Distance from pin {index - 1}: {km:.0f} km."
420
- return line
421
-
422
-
423
- def _draw_land(ax, linewidth: float) -> None:
424
- for feature in _countries():
425
- for ring in _rings(feature["geometry"]):
426
- ax.add_patch(
427
- MplPolygon(
428
- ring,
429
- closed=True,
430
- facecolor=_LAND,
431
- edgecolor=_BORDER,
432
- linewidth=linewidth,
433
- )
434
- )
435
-
436
-
437
- def _label_countries(ax, lat: float, lon: float, span: float) -> None:
438
- """Label countries visible in the window, clipped to the viewport."""
439
- for feature in _countries():
440
- for ring in _rings(feature["geometry"]):
441
- visible = [
442
- p for p in ring if abs(p[0] - lon) < span and abs(p[1] - lat) < span
443
- ]
444
- if len(visible) > 3:
445
- cx = sum(p[0] for p in visible) / len(visible)
446
- cy = sum(p[1] for p in visible) / len(visible)
447
- ax.text(
448
- cx,
449
- cy,
450
- feature["properties"]["NAME"],
451
- fontsize=7.4,
452
- ha="center",
453
- color="#55504a",
454
- zorder=7,
455
- )
456
- break
457
-
458
-
459
- def _draw_detail(ax, lat: float, lon: float, span: float) -> None:
460
- """Draw urban areas, rivers, roads and town names, by zoom level.
461
-
462
- Each layer switches on below its own span so a wide view stays legible and a
463
- tight view shows enough road structure to aim a pin within a town.
464
- """
465
- if span <= URBAN_MAX_SPAN_DEG:
466
- for row in _detail("urban"):
467
- for part in _parts(row):
468
- if _near(part, lat, lon, span):
469
- ax.add_patch(
470
- MplPolygon(
471
- part,
472
- closed=True,
473
- facecolor=_URBAN,
474
- edgecolor="none",
475
- zorder=2,
476
- )
477
- )
478
- if span <= RIVERS_MAX_SPAN_DEG:
479
- segments = [
480
- part
481
- for row in _detail("rivers")
482
- for part in _parts(row)
483
- if _near(part, lat, lon, span)
484
- ]
485
- if segments:
486
- ax.add_collection(
487
- LineCollection(segments, colors=_RIVER, linewidths=0.9, zorder=3)
488
- )
489
- if span <= ROADS_MAX_SPAN_DEG:
490
- segments = [
491
- part
492
- for row in _detail("roads")
493
- for part in _parts(row)
494
- if _near(part, lat, lon, span)
495
- ]
496
- if segments:
497
- ax.add_collection(
498
- LineCollection(segments, colors=_ROAD, linewidths=1.4, zorder=4)
499
- )
500
- if span <= STREET_MAX_SPAN_DEG and _STREETS_ENABLED[0]:
501
- ways = street_ways(lat, lon, span)
502
- # Generalise by zoom the way a real style does: drawing every service
503
- # road and footpath at city scale turns the grid into a smear, so the
504
- # minor classes only appear once the window is tight enough to hold
505
- # them, and widths grow as the window shrinks.
506
- floor = 0.75 if span > 0.15 else (0.55 if span > 0.05 else 0.0)
507
- scale = 1.0 if span > 0.15 else (1.4 if span > 0.05 else 1.9)
508
- drawn = [way for way in ways if way["w"] >= floor]
509
- for weight in sorted({way["w"] for way in drawn}):
510
- segments = [way["c"] for way in drawn if way["w"] == weight]
511
- width = weight * scale
512
- # A casing under a white fill is what makes a dense grid legible; a
513
- # single flat colour reads as noise.
514
- ax.add_collection(
515
- LineCollection(
516
- segments, colors="#c9bfb3", linewidths=width + 0.5, zorder=4
517
- )
518
- )
519
- ax.add_collection(
520
- LineCollection(segments, colors="#ffffff", linewidths=width, zorder=5)
521
- )
522
- _label_streets(ax, drawn, lat, lon, span)
523
- if span <= TOWNS_MAX_SPAN_DEG:
524
- shown = 0
525
- for row in _detail("places"):
526
- point = row["c"]
527
- if abs(point[0] - lon) < span * 0.95 and abs(point[1] - lat) < span * 0.95:
528
- ax.plot(
529
- point[0],
530
- point[1],
531
- "o",
532
- markersize=2.6,
533
- markerfacecolor=_TOWN,
534
- markeredgecolor="none",
535
- zorder=6,
536
- )
537
- if row.get("n"):
538
- label = ax.text(
539
- point[0],
540
- point[1] + span * 0.03,
541
- row["n"],
542
- fontsize=6.2,
543
- ha="center",
544
- color=_TOWN,
545
- zorder=9,
546
- )
547
- label.set_path_effects(
548
- [
549
- path_effects.Stroke(linewidth=1.8, foreground="#ffffff"),
550
- path_effects.Normal(),
551
- ]
552
- )
553
- shown += 1
554
- if shown >= (8 if span <= STREET_MAX_SPAN_DEG else 28):
555
- break
556
-
557
-
558
- def _label_streets(ax, ways: list[dict], lat: float, lon: float, span: float) -> None:
559
- """
560
- Write street names along the ways, the way a real map style does.
561
-
562
- One label per name, on that name's longest visible run, rotated to follow
563
- the road and haloed so it stays readable over the casing. Longest-run
564
- selection matters: labelling an arbitrary segment puts "Main Street" on a
565
- 50 m stub while the avenue itself goes unnamed.
566
-
567
- Args:
568
- ax:
569
- Matplotlib axes to draw on.
570
- ways (`list[dict]`):
571
- Way records from [`street_ways`], some carrying a name in `n`.
572
- lat (`float`):
573
- Latitude at the centre of the window.
574
- lon (`float`):
575
- Longitude at the centre of the window.
576
- span (`float`):
577
- Half-width of the window in degrees.
578
- """
579
- if span > STREET_LABEL_MAX_SPAN_DEG:
580
- return
581
- # Pick, per name, the longest run of points that actually falls inside the
582
- # window, so the label lands where the reader can see it.
583
- best: dict[str, tuple[float, list]] = {}
584
- for way in ways:
585
- name = way.get("n")
586
- if not name:
587
- continue
588
- inside = [
589
- point
590
- for point in way["c"]
591
- if abs(point[0] - lon) < span * 0.92 and abs(point[1] - lat) < span * 0.92
592
- ]
593
- if len(inside) < 2:
594
- continue
595
- length = sum(
596
- math.dist(inside[i], inside[i + 1]) for i in range(len(inside) - 1)
597
- )
598
- if name not in best or length > best[name][0]:
599
- best[name] = (length, inside)
600
-
601
- ordered = sorted(best.items(), key=lambda item: -item[1][0])
602
- aspect = max(0.05, math.cos(math.radians(lat)))
603
- # Parallel streets in a grid all have their midpoint in the same place, so
604
- # placing every label at its midpoint stacks them into an unreadable pile.
605
- # Claim one coarse cell per label, walking along each road to find a free
606
- # one, and drop the label rather than overprint. Longest roads go first, so
607
- # the ones worth naming win the space.
608
- occupied: set[tuple[int, int]] = set()
609
- # The cell has to be taller than the crowding you want to break up, not
610
- # just taller than the glyphs: parallel streets one block apart land in
611
- # different fine cells and still read as a stack. A tall cell forces
612
- # neighbours to slide along their own road instead, which is what a real
613
- # map style does.
614
- cell_x = span * 0.22
615
- cell_y = span * 0.13
616
- placed = 0
617
- for name, (_, points) in ordered:
618
- if placed >= MAX_STREET_LABELS:
619
- break
620
- middle = len(points) // 2
621
- # Try the midpoint first, then positions either side of it.
622
- order = sorted(range(len(points)), key=lambda i: abs(i - middle))
623
- chosen = None
624
- for i in order:
625
- cell = (
626
- int((points[i][0] - lon) / cell_x),
627
- int((points[i][1] - lat) / cell_y),
628
- )
629
- if cell not in occupied:
630
- occupied.add(cell)
631
- chosen = i
632
- break
633
- if chosen is None:
634
- continue
635
- placed += 1
636
- middle = chosen
637
- start = points[max(0, middle - 1)]
638
- end = points[min(len(points) - 1, middle + 1)]
639
- # Longitude degrees are shorter than latitude ones away from the
640
- # equator, so the on-screen angle needs the cos(lat) correction or
641
- # labels sit visibly off their road.
642
- angle = math.degrees(
643
- math.atan2(end[1] - start[1], (end[0] - start[0]) * aspect)
644
- )
645
- if angle > 90:
646
- angle -= 180
647
- elif angle < -90:
648
- angle += 180
649
- text = ax.text(
650
- points[middle][0],
651
- points[middle][1],
652
- name,
653
- fontsize=4.6,
654
- ha="center",
655
- va="center",
656
- rotation=angle,
657
- rotation_mode="anchor",
658
- color="#4a453f",
659
- zorder=8,
660
- )
661
- text.set_path_effects(
662
- [
663
- path_effects.Stroke(linewidth=1.4, foreground="#ffffff"),
664
- path_effects.Normal(),
665
- ]
666
- )
667
-
668
-
669
- def _scale_bar(ax, lat: float, lon: float, span: float) -> None:
670
- km_per_degree = 111.32 * max(0.05, math.cos(math.radians(lat)))
671
- # A bar reading "100 km" across a hemisphere is worse than no bar.
672
- if span > 20:
673
- unit = 2000
674
- elif span > 5:
675
- unit = 500
676
- elif span > 2:
677
- unit = 100
678
- elif span > 0.5:
679
- unit = 20
680
- else:
681
- unit = 5
682
- bar = unit / km_per_degree
683
- x0, y0 = lon - span * 0.9, lat - span * 0.9
684
- ax.plot([x0, x0 + bar], [y0, y0], color=_INK, linewidth=2.2, zorder=9)
685
- ax.text(
686
- x0 + bar / 2,
687
- y0 + span * 0.04,
688
- f"{unit} km",
689
- fontsize=6.6,
690
- ha="center",
691
- color=_INK,
692
- zorder=9,
693
- )
694
-
695
-
696
- def _plot_pins(ax, pins: list[tuple[float, float]], size: float) -> None:
697
- for i, (lat, lon) in enumerate(pins, 1):
698
- ax.plot(
699
- lon,
700
- lat,
701
- marker="o",
702
- markersize=size,
703
- markerfacecolor=_PIN,
704
- markeredgecolor="white",
705
- markeredgewidth=1.5,
706
- zorder=8,
707
- )
708
- ax.annotate(
709
- str(i),
710
- (lon, lat),
711
- color="white",
712
- fontsize=size * 0.62,
713
- weight="bold",
714
- ha="center",
715
- va="center",
716
- zorder=9,
717
- )
718
-
719
-
720
- def _plot_truth(ax, pins, truth: tuple[float, float], size: float) -> None:
721
- """Draw the true location and the line to the guess it is being compared to."""
722
- truth_lat, truth_lon = truth
723
- if pins:
724
- guess_lat, guess_lon = pins[-1]
725
- ax.plot(
726
- [guess_lon, truth_lon],
727
- [guess_lat, truth_lat],
728
- linestyle="--",
729
- color=_PIN,
730
- linewidth=1.6,
731
- zorder=7,
732
- )
733
- ax.plot(
734
- truth_lon,
735
- truth_lat,
736
- marker="o",
737
- markersize=size,
738
- markerfacecolor=_TRUTH,
739
- markeredgecolor="white",
740
- markeredgewidth=1.5,
741
- zorder=10,
742
- )
743
-
744
-
745
- def render_map(
746
- pins: list[tuple[float, float]],
747
- focus: tuple[float, float] | None = None,
748
- span_deg: float = 7.0,
749
- dpi: int = 100,
750
- truth: tuple[float, float] | None = None,
751
- ) -> Image.Image:
752
- """
753
- Render the guess map: a world panel plus a zoomed panel.
754
-
755
- Args:
756
- pins (`list[tuple[float, float]]`):
757
- Pins as `(lat, lon)`, drawn and numbered in order.
758
- focus (`tuple[float, float]`, *optional*):
759
- Centre of the zoomed panel. Defaults to the last pin.
760
- span_deg (`float`, *optional*, defaults to `7.0`):
761
- Half-width of the zoomed panel in degrees. Below roughly 4 degrees
762
- the panel adds urban areas, roads, rivers and town names, when the
763
- optional detail layers are installed.
764
- dpi (`int`, *optional*, defaults to `100`):
765
- Figure resolution.
766
- truth (`tuple[float, float]`, *optional*):
767
- Ground truth, drawn in green with a dashed line to the last pin.
768
- Only ever passed after a guess has been scored, so it cannot leak
769
- into an observation the agent sees before committing.
770
-
771
- Returns:
772
- `PIL.Image.Image`: The two-panel map.
773
- """
774
- fig, (world, zoom) = plt.subplots(
775
- 1, 2, figsize=(10.2, 3.9), dpi=dpi, gridspec_kw={"width_ratios": [1.55, 1]}
776
- )
777
-
778
- _draw_land(world, 0.35)
779
- world.set_xlim(-180, 180)
780
- world.set_ylim(-90, 90)
781
- world.set_facecolor(_WATER)
782
- for x in range(-180, 181, 60):
783
- world.axvline(x, color="white", linewidth=0.5, alpha=0.7)
784
- for y in range(-60, 61, 30):
785
- world.axhline(y, color="white", linewidth=0.5, alpha=0.7)
786
- _plot_pins(world, pins, 8.0)
787
- if truth is not None:
788
- _plot_truth(world, pins, truth, 8.0)
789
- title = "guess and true location" if truth is not None else "your pins - world"
790
- world.set_title(title, fontsize=9, loc="left", color="#333")
791
- world.set_xticks([])
792
- world.set_yticks([])
793
-
794
- focus_lat, focus_lon = focus if focus else (pins[-1] if pins else (20.0, 0.0))
795
- _draw_land(zoom, 0.9)
796
- zoom.set_xlim(focus_lon - span_deg, focus_lon + span_deg)
797
- zoom.set_ylim(focus_lat - span_deg, focus_lat + span_deg)
798
- zoom.set_facecolor(_WATER)
799
- _draw_detail(zoom, focus_lat, focus_lon, span_deg)
800
- # Above about 20 degrees every country in a hemisphere wants a label and the
801
- # panel turns into a stack of overlapping text.
802
- if TOWNS_MAX_SPAN_DEG < span_deg <= 20.0:
803
- _label_countries(zoom, focus_lat, focus_lon, span_deg)
804
- _plot_pins(zoom, pins, 10.0)
805
- if truth is not None:
806
- _plot_truth(zoom, pins, truth, 10.0)
807
- _scale_bar(zoom, focus_lat, focus_lon, span_deg)
808
- # Below a degree, degrees round to "0"; kilometres are the useful unit there.
809
- width_deg = span_deg * 2
810
- label = (
811
- f"{width_deg:.0f} deg view"
812
- if width_deg >= 1
813
- else f"{width_deg * 111:.0f} km view"
814
- )
815
- zoom.set_title(label, fontsize=9, loc="left", color="#333")
816
- zoom.set_xticks([])
817
- zoom.set_yticks([])
818
-
819
- fig.tight_layout(pad=0.6)
820
- buf = io.BytesIO()
821
- fig.savefig(buf, format="png", facecolor="white")
822
- plt.close(fig)
823
- buf.seek(0)
824
- return Image.open(buf).convert("RGB")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/render/pano.py DELETED
@@ -1,108 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Perspective views out of an equirectangular panorama.
4
-
5
- `look()` is a gnomonic reprojection: build a camera ray for every output
6
- pixel, rotate it by the requested heading and pitch, convert to spherical
7
- coordinates and sample the source image. Integer sampling keeps it
8
- deterministic — the same arguments always produce the same bytes, which is
9
- what lets a GRPO group share one starting observation.
10
- """
11
-
12
- from __future__ import annotations
13
-
14
- import base64
15
- import io
16
-
17
- import numpy as np
18
- from PIL import Image
19
-
20
-
21
- DEFAULT_SIZE = (640, 640)
22
-
23
-
24
- def look(
25
- pano: Image.Image,
26
- heading_deg: float,
27
- pitch_deg: float = 0.0,
28
- fov_deg: float = 90.0,
29
- size: tuple[int, int] = DEFAULT_SIZE,
30
- ) -> Image.Image:
31
- """
32
- Render one perspective view out of an equirectangular panorama.
33
-
34
- Args:
35
- pano (`PIL.Image.Image`):
36
- Source panorama, 2:1 equirectangular.
37
- heading_deg (`float`):
38
- Compass heading in degrees, `0` being the panorama's own north.
39
- pitch_deg (`float`, *optional*, defaults to `0.0`):
40
- Vertical angle in degrees; positive looks up.
41
- fov_deg (`float`, *optional*, defaults to `90.0`):
42
- Horizontal field of view. Smaller values zoom in.
43
- size (`tuple[int, int]`, *optional*, defaults to `(640, 640)`):
44
- Output width and height in pixels.
45
-
46
- Returns:
47
- `PIL.Image.Image`: The rendered view.
48
-
49
- Examples:
50
-
51
- ```python
52
- view = look(Image.open("pano.jpg"), heading_deg=90, fov_deg=30)
53
- ```
54
- """
55
- src = np.asarray(pano.convert("RGB"))
56
- src_h, src_w = src.shape[:2]
57
- width, height = size
58
-
59
- focal = 0.5 * width / np.tan(np.radians(fov_deg) / 2)
60
- xs, ys = np.meshgrid(np.arange(width) - width / 2, np.arange(height) - height / 2)
61
- rays = np.stack([xs, -ys, np.full_like(xs, focal, dtype=float)], axis=-1)
62
- rays /= np.linalg.norm(rays, axis=-1, keepdims=True)
63
-
64
- pitch, heading = np.radians(pitch_deg), np.radians(heading_deg)
65
- rot_x = np.array(
66
- [
67
- [1, 0, 0],
68
- [0, np.cos(pitch), -np.sin(pitch)],
69
- [0, np.sin(pitch), np.cos(pitch)],
70
- ]
71
- )
72
- rot_y = np.array(
73
- [
74
- [np.cos(heading), 0, np.sin(heading)],
75
- [0, 1, 0],
76
- [-np.sin(heading), 0, np.cos(heading)],
77
- ]
78
- )
79
- rays = rays @ rot_x.T @ rot_y.T
80
-
81
- lon = np.arctan2(rays[..., 0], rays[..., 2])
82
- lat = np.arcsin(np.clip(rays[..., 1], -1.0, 1.0))
83
- u = ((lon / (2 * np.pi) + 0.5) * src_w).astype(np.int32) % src_w
84
- v = np.clip(((0.5 - lat / np.pi) * src_h).astype(np.int32), 0, src_h - 1)
85
- return Image.fromarray(src[v, u])
86
-
87
-
88
- def to_base64(image: Image.Image, fmt: str = "JPEG", quality: int = 85) -> str:
89
- """
90
- Encode an image for transport inside an observation.
91
-
92
- Args:
93
- image (`PIL.Image.Image`):
94
- Image to encode.
95
- fmt (`str`, *optional*, defaults to `"JPEG"`):
96
- Pillow format name. Views use JPEG; maps use PNG.
97
- quality (`int`, *optional*, defaults to `85`):
98
- JPEG quality, ignored for PNG.
99
-
100
- Returns:
101
- `str`: Base64-encoded image bytes, without a data URI prefix.
102
- """
103
- buf = io.BytesIO()
104
- if fmt.upper() == "JPEG":
105
- image.save(buf, format="JPEG", quality=quality, optimize=True)
106
- else:
107
- image.save(buf, format=fmt)
108
- return base64.b64encode(buf.getvalue()).decode("ascii")
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/build/lib/geoguesser_env/server/scoring.py DELETED
@@ -1,233 +0,0 @@
1
- # SPDX-License-Identifier: BSD-3-Clause
2
-
3
- """Distance scoring and action costs.
4
-
5
- The distance curve is GeoGuessr's own, `5000 * exp(-d / 1492.7)`, normalised
6
- to `[0, 1]`. Keeping the real curve means scores are directly interpretable
7
- against the game most people already know.
8
- """
9
-
10
- from __future__ import annotations
11
-
12
- import math
13
-
14
-
15
- # GeoGuessr's decay constant, in kilometres.
16
- DECAY_KM = 1492.7
17
-
18
- # A second, much slower decay used only by the `"mixture"` reward shape. The
19
- # game curve is worth 0.018 across the whole 6000-20000 km range, so a policy
20
- # that lands on the wrong continent -- a third of rollouts for a 4B model --
21
- # gets no gradient for getting *less* wrong. This scale restores one.
22
- LONG_DECAY_KM = 5000.0
23
-
24
- # Fraction of the mixture carried by the game curve; the rest is the long scale.
25
- MIXTURE_SHORT_WEIGHT = 0.5
26
-
27
- REWARD_SHAPES = ("geoguessr", "mixture")
28
-
29
- EARTH_RADIUS_KM = 6371.0088
30
-
31
- # Action costs. Information gathering is cheap but not free, so an episode has
32
- # to trade breadth of search against committing to a guess.
33
- COST_LOOK = 0.01
34
- COST_MAP = 0.01
35
- COST_PIN = 0.02
36
- COST_MOVE = 0.05
37
-
38
- # Partial credit weights, used when the reward is hierarchical.
39
- WEIGHT_COUNTRY = 0.15
40
- WEIGHT_REGION = 0.10
41
-
42
-
43
- def haversine_km(lat_a: float, lon_a: float, lat_b: float, lon_b: float) -> float:
44
- """
45
- Great-circle distance between two points in kilometres.
46
-
47
- Args:
48
- lat_a (`float`):
49
- Latitude of the first point in degrees.
50
- lon_a (`float`):
51
- Longitude of the first point in degrees.
52
- lat_b (`float`):
53
- Latitude of the second point in degrees.
54
- lon_b (`float`):
55
- Longitude of the second point in degrees.
56
-
57
- Returns:
58
- `float`: Distance in kilometres.
59
-
60
- Examples:
61
-
62
- ```python
63
- d = haversine_km(-16.4897, -68.1193, -17.7833, -63.1821)
64
- ```
65
- """
66
- phi_a, phi_b = math.radians(lat_a), math.radians(lat_b)
67
- d_phi = phi_b - phi_a
68
- d_lambda = math.radians(lon_b - lon_a)
69
- h = (
70
- math.sin(d_phi / 2) ** 2
71
- + math.cos(phi_a) * math.cos(phi_b) * math.sin(d_lambda / 2) ** 2
72
- )
73
- return 2 * EARTH_RADIUS_KM * math.asin(math.sqrt(min(1.0, h)))
74
-
75
-
76
- def distance_score(distance_km: float, shape: str = "geoguessr") -> float:
77
- """
78
- Map a distance to a score in `[0, 1]`.
79
-
80
- Args:
81
- distance_km (`float`):
82
- Distance between guess and truth in kilometres.
83
- shape (`str`, *optional*, defaults to `"geoguessr"`):
84
- `"geoguessr"` for the game's own curve, or `"mixture"` for the
85
- two-scale curve used when training. The mixture is deliberately
86
- *not* the default: reported scores stay comparable to the game.
87
-
88
- Returns:
89
- `float`: For `"geoguessr"`, `exp(-distance_km / 1492.7)` -- 0 km scores
90
- `1.0`, 150 km about `0.90`, 5000 km about `0.035`. For `"mixture"`,
91
- half that plus half of `exp(-distance_km / 5000)`, which keeps the
92
- mid-range sharp while leaving real gradient past 3000 km.
93
-
94
- Examples:
95
-
96
- ```python
97
- near = distance_score(200.0) # 0.875
98
- far = distance_score(8000.0, shape="mixture") # 0.104, versus 0.005
99
- ```
100
- """
101
- if distance_km < 0:
102
- raise ValueError(f"distance_km must be non-negative, got {distance_km}")
103
- if shape not in REWARD_SHAPES:
104
- raise ValueError(f"shape must be one of {REWARD_SHAPES}, got {shape!r}")
105
- short = math.exp(-distance_km / DECAY_KM)
106
- if shape == "geoguessr":
107
- return short
108
- long = math.exp(-distance_km / LONG_DECAY_KM)
109
- return MIXTURE_SHORT_WEIGHT * short + (1.0 - MIXTURE_SHORT_WEIGHT) * long
110
-
111
-
112
- def action_cost(
113
- n_looks: int = 0, n_maps: int = 0, n_pins: int = 0, n_moves: int = 0
114
- ) -> float:
115
- """
116
- Total cost of the information gathering done this episode.
117
-
118
- Args:
119
- n_looks (`int`, *optional*, defaults to `0`):
120
- Number of view renders.
121
- n_maps (`int`, *optional*, defaults to `0`):
122
- Number of map views that were not pins.
123
- n_pins (`int`, *optional*, defaults to `0`):
124
- Number of pins placed.
125
- n_moves (`int`, *optional*, defaults to `0`):
126
- Number of moves taken.
127
-
128
- Returns:
129
- `float`: Cost to subtract from the distance score.
130
- """
131
- return (
132
- COST_LOOK * n_looks
133
- + COST_MAP * n_maps
134
- + COST_PIN * n_pins
135
- + COST_MOVE * n_moves
136
- )
137
-
138
-
139
- # The multiplicative cost is capped so a very long episode scales the score down
140
- # rather than erasing it. Without a cap a 20-move episode would reach zero.
141
- MAX_COST_FRACTION = 0.5
142
-
143
-
144
- def compute_reward(
145
- distance_km: float | None,
146
- *,
147
- cost: float = 0.0,
148
- country_hit: bool = False,
149
- region_hit: bool = False,
150
- hierarchical: bool = False,
151
- shape: str = "geoguessr",
152
- cost_mode: str = "subtract",
153
- ) -> float:
154
- """
155
- Combine distance, partial credit and action cost into one reward.
156
-
157
- Args:
158
- distance_km (`float` or `None`):
159
- Distance from guess to truth. `None` means the guess could not be
160
- parsed, which scores zero before costs.
161
- cost (`float`, *optional*, defaults to `0.0`):
162
- Accumulated action cost.
163
- country_hit (`bool`, *optional*, defaults to `False`):
164
- Whether the guessed country matched.
165
- region_hit (`bool`, *optional*, defaults to `False`):
166
- Whether the guessed region matched.
167
- hierarchical (`bool`, *optional*, defaults to `False`):
168
- Whether to add country and region partial credit.
169
- shape (`str`, *optional*, defaults to `"geoguessr"`):
170
- Distance curve, passed to [`~scoring.distance_score`].
171
- cost_mode (`str`, *optional*, defaults to `"subtract"`):
172
- `"subtract"` reproduces the game: the cost comes off the score and
173
- the result is floored at zero. `"multiply"` scales the score by
174
- `1 - cost` instead, which is what training wants -- see below.
175
-
176
- Returns:
177
- `float`: Reward in `[0, 1]`.
178
-
179
- <Tip warning={true}>
180
-
181
- `"subtract"` and a floor at zero destroy the ordering of bad guesses. Mean
182
- cost for a 4B model is 0.13, and the game curve falls below that at about
183
- 3300 km, so a 3324 km miss and an 18723 km miss both score exactly 0.0 --
184
- measured across 200 episodes, 77 of them collapsed to a single value with
185
- zero variance. A GRPO group drawn from those has no advantage and therefore
186
- contributes no gradient. `"multiply"` cannot do this: scaling by a positive
187
- factor preserves the ordering whatever the cost.
188
-
189
- </Tip>
190
-
191
- Examples:
192
-
193
- ```python
194
- # Training: gradient survives on the wrong continent.
195
- reward = compute_reward(8000.0, cost=0.13, shape="mixture", cost_mode="multiply")
196
- ```
197
- """
198
- if distance_km is None:
199
- return 0.0
200
- score = distance_score(distance_km, shape=shape)
201
- if hierarchical:
202
- score += WEIGHT_COUNTRY * country_hit + WEIGHT_REGION * region_hit
203
- score = min(1.0, score)
204
- if cost_mode == "multiply":
205
- return score * (1.0 - min(max(cost, 0.0), MAX_COST_FRACTION))
206
- if cost_mode != "subtract":
207
- raise ValueError(
208
- f"cost_mode must be 'subtract' or 'multiply', got {cost_mode!r}"
209
- )
210
- return max(0.0, score - cost)
211
-
212
-
213
- def verdict(distance_km: float) -> str:
214
- """
215
- Short human-readable label for a distance, for UIs and logs.
216
-
217
- Args:
218
- distance_km (`float`):
219
- Distance between guess and truth in kilometres.
220
-
221
- Returns:
222
- `str`: One of `"Perfect"`, `"Pinpoint"`, `"Close"`, `"Right region"`
223
- or `"Wrong continent"`.
224
- """
225
- if distance_km < 0.025:
226
- return "Perfect"
227
- if distance_km < 25:
228
- return "Pinpoint"
229
- if distance_km < 200:
230
- return "Close"
231
- if distance_km < 1500:
232
- return "Right region"
233
- return "Wrong continent"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/openenv_geoguesser_env.egg-info/PKG-INFO DELETED
@@ -1,760 +0,0 @@
1
- Metadata-Version: 2.4
2
- Name: openenv-geoguesser-env
3
- Version: 0.1.0
4
- Summary: GeoGuessr-style visual geolocation environment for OpenEnv
5
- Requires-Python: >=3.10
6
- Description-Content-Type: text/markdown
7
- Requires-Dist: openenv>=0.3.1
8
- Requires-Dist: fastapi>=0.115.0
9
- Requires-Dist: pydantic>=2.0.0
10
- Requires-Dist: uvicorn>=0.24.0
11
- Requires-Dist: fastmcp>=2.0.0
12
- Requires-Dist: pillow>=10.0.0
13
- Requires-Dist: numpy>=1.24.0
14
- Requires-Dist: matplotlib>=3.7.0
15
- Provides-Extra: ui
16
- Requires-Dist: gradio>=4.0.0; extra == "ui"
17
- Provides-Extra: dev
18
- Requires-Dist: pytest>=8.0.0; extra == "dev"
19
-
20
- # GeoGuesser
21
-
22
- A GeoGuessr-style visual geolocation environment. The agent is dropped at an
23
- unknown street-level location, looks around, walks along the road, pins
24
- candidate coordinates on a map to check itself, and commits to a final guess.
25
- Reward is distance-based, using the game's own scoring curve.
26
-
27
- Independent open-source project, unaffiliated with GeoGuessr AB. Imagery comes
28
- from Mapillary contributors under CC-BY-SA-4.0.
29
-
30
- ## Quick start
31
-
32
- ```bash
33
- cd envs/geoguesser_env
34
-
35
- # The frozen 200-task eval split is committed, so this runs as-is, with the
36
- # same configuration the Space uses.
37
- ./scripts/serve_local.sh # http://localhost:8000/web/
38
-
39
- # To build your own data (needs a free Mapillary token with READ scope):
40
- export MAPILLARY_API_KEY_TRAIN="MLY|..."
41
- python scripts/harvest_tiles.py # enumerate sequences
42
- ./scripts/build_dataset.sh # mirror tasks offline
43
- python scripts/verify_offline.py tasks/pool_offline_5k.jsonl
44
- python scripts/split_tasks.py tasks/pool_offline_5k.jsonl --eval 200
45
- ```
46
-
47
- ```python
48
- from geoguesser_env import GeoGuesserEnv, GuessAction, LookAction, PinAction
49
-
50
- env = GeoGuesserEnv(base_url="http://localhost:8000")
51
-
52
- result = env.reset(split="eval", index=7) # byte-identical on repeat
53
- print(result.observation.prompt)
54
-
55
- result = env.step(LookAction(heading_deg=90, fov_deg=45))
56
- result = env.step(PinAction(lat=-16.5, lon=-68.1))
57
- print(result.observation.feedback)
58
- # Pin 1 placed at -16.5000, -68.1000 - Bolivia (South America).
59
- # Nearest major city: La Paz, ~5 km E. 10 actions left.
60
-
61
- result = env.step(GuessAction(response="Altiplano. <guess>-16.49, -68.12</guess>"))
62
- print(result.reward, result.observation.distance_km)
63
- ```
64
-
65
- ## Tools
66
-
67
- | Tool | What it does | Cost |
68
- |------|--------------|------|
69
- | `look(heading_deg, pitch_deg, fov_deg)` | Render a view. Heading is absolute, `0` is true north | −0.01 |
70
- | `pan(delta_deg)` | Turn relative to the current heading | −0.01 |
71
- | `zoom(fov_deg)` | Narrow the field of view; around 30 reads distant signs | −0.01 |
72
- | `move(direction, meters)` | Walk the captured road; reports distance actually travelled | −0.05 |
73
- | `place_pin(lat, lon, label)` | Pin a candidate and see where it falls on the map | −0.02 |
74
- | `view_map(lat, lon, span_deg)` | Pan and zoom the map without pinning | −0.01 |
75
- | `list_pins()` / `clear_pins()` | Review or drop candidates | free |
76
- | `measure(lat_a, lon_a, lat_b, lon_b)` | Distance between two of your own points | free |
77
- | `reverse_geocode(lat, lon)` | Name the country and nearest city at a coordinate | free |
78
- | `submit_guess(lat, lon, ...)` | Commit the answer. Terminal | — |
79
-
80
- Tools the backend cannot serve are **not registered**, so the agent never sees
81
- a tool that always fails.
82
-
83
- ### The agent's map and the player's map agree
84
-
85
- The player sees live OpenFreeMap tiles; the agent sees an offline Natural Earth
86
- render. They have to agree about *how precisely a pin can be aimed*, because
87
- that is what the distance reward measures — a map showing only country outlines
88
- lets you place a country, not a point within a city.
89
-
90
- So the guess map is zoom-aware. `place_pin` takes `span_deg`, and the render
91
- adds detail as the window tightens:
92
-
93
- | Window | What the agent's map shows |
94
- |--------|----------------------------|
95
- | wider than ~4 deg | coastlines, borders, country names |
96
- | under ~4 deg | urban areas, highways, rivers, town names (Natural Earth 10m) |
97
- | under ~0.35 deg | **real OSM streets**, fetched from Overpass and cached |
98
-
99
- Natural Earth tops out at highway level — it shows the motorways around a city
100
- but not the grid inside it. Below 0.35 degrees the map therefore fetches actual
101
- ways from Overpass, generalising by zoom the way a real style does: minor
102
- classes appear only once the window is tight enough to hold them, and widths
103
- grow as it shrinks. A pin on Abuja at `span_deg=0.05` came back as an 11 km
104
- window with the full street grid, drawn white-on-pale to read like the
105
- player's Positron tiles.
106
-
107
- Overpass has real limits, and they are the binding constraint on how this
108
- scales: roughly **10,000 requests and 1 GB per day**, about **2 concurrent
109
- slots per IP**, a 180 s runtime and 512 MiB memory ceiling per query, HTTP 429
110
- when rate limited and 504 when a query is too large. Cooldowns lengthen for
111
- heavy users. So street detail is right for eval, demos and modest training, and
112
- the cache is what keeps it polite — a run doing millions of pins must pre-warm
113
- or bundle a Protomaps extract instead.
114
-
115
- The first render of a neighbourhood costs 3-16 s; every later one is served
116
- from `data/geo/osm_cache/` in ~30 ms and is byte-identical. That makes an
117
- episode deterministic once warm, and a frozen eval should pre-warm the cache
118
- the same way it pre-warms panoramas — or set
119
- `GEOGUESSER_STREET_DETAIL=0`, which falls back to Natural Earth and never
120
- touches the network. Any fetch failure degrades to no streets rather than
121
- failing the step.
122
-
123
- The optional detail layers are fetched once, since 87 MB of GeoJSON does not
124
- belong in the repo:
125
-
126
- ```bash
127
- python scripts/fetch_detail_geo.py # compacts to ~39 MB, gitignored
128
- ```
129
-
130
- Without them the map still renders, with outlines and major cities only.
131
-
132
- In the play page the pin carries the zoom you are actually looking at, so the
133
- "what the agent sees" panel is framed like your own view at the same scale and
134
- with comparable detail.
135
-
136
- Overpass has a usage policy that discourages heavy automated querying, so this
137
- is right for eval, demos and modest training, and the cache is what keeps it
138
- polite. A run doing millions of pins should pre-warm or bundle a Protomaps
139
- extract instead.
140
-
141
- ### Zoom resolves real detail
142
-
143
- Zooming is not cosmetic, but it needs the right source. A 30-degree view of a
144
- 2048x1024 panorama samples only about 170 source pixels, so narrowing the field
145
- of view barely adds information — measured mean gradient 6.60 at 90 degrees
146
- against 7.03 at 30. The 7680x3840 original roughly doubles it (10.23 against
147
- 14.87), which is the difference between guessing at a sign and reading it.
148
-
149
- So each panorama is cached twice. Wide views render from the 2048 derivative in
150
- ~30 ms; a field of view at or below 45 degrees pulls the original and renders in
151
- ~70 ms. If no original exists the step degrades to a soft view rather than
152
- failing. Set `GEOGUESSER_HIRES_ZOOM=0` to disable it.
153
-
154
- ### Pinning tells you where you pointed, not whether you are right
155
-
156
- `place_pin` returns a rendered map and a description of the pinned location:
157
- country, subregion, nearest city with distance and bearing, and the distance
158
- to the agent's own earlier pins. It reveals nothing about the target.
159
-
160
- That restraint is deliberate. Any signal about the truth — a distance, a
161
- warmer/colder hint — would make binary search the optimal policy, and the
162
- environment would measure bisection rather than geographic reasoning. Distance
163
- and score arrive only from `submit_guess`.
164
-
165
- ## Reward
166
-
167
- ```
168
- geo = exp(-distance_km / 1492.7) # GeoGuessr's curve, in [0, 1]
169
- partial = 0.15 * country_hit + 0.10 * region_hit # when hierarchical
170
- cost = 0.01*looks + 0.01*maps + 0.02*pins + 0.05*moves
171
- reward = clip(geo + partial, 0, 1) - cost
172
- ```
173
-
174
- An unparseable or out-of-range guess scores `0.0` and says why. Parsing
175
- accepts what models actually emit: decimal pairs, DMS (`48°51'29"N`), labelled
176
- `lat:`/`lon:`, JSON, and `<guess>` tags.
177
-
178
- ### Training against it
179
-
180
- The defaults above reproduce the game, which makes a score directly comparable
181
- to GeoGuessr. They are the wrong shape for RL, and three flags change that:
182
-
183
- | flag | play / eval | training | why |
184
- |---|---|---|---|
185
- | `reward_shape` | `"geoguessr"` | `"mixture"` | The game curve is worth **0.018** across the whole 6000-20000 km range, so a policy gets no gradient for landing on the right continent instead of the wrong one. `"mixture"` adds a 5000 km scale, making that span worth 0.150. |
186
- | `cost_mode` | `"subtract"` | `"multiply"` | Mean action cost for a 4B model is 0.13 and the curve falls below that at ~3300 km, so `max(0, geo - cost)` floors every worse guess at exactly zero. Measured over 200 episodes: **77 collapsed to 0.0 with zero variance**, so a GRPO group drawn from them has no advantage and yields no gradient. A multiplier cannot do this. |
187
- | `hide_task_identity` | `False` | **`True`** | `metadata` carries `attribution.creator_username`, and the contributor determines the country outright for **74% of training tasks** (`amsterdam` only maps the Netherlands). `task_index`/`task_id`/`sequence_id` are a few thousand memorisable keys straight to a coordinate. Either lets a policy score without reading the image. |
188
-
189
- ```bash
190
- GEOGUESSER_REWARD_SHAPE=mixture \
191
- GEOGUESSER_COST_MODE=multiply \
192
- GEOGUESSER_HIDE_IDENTITY=1 \
193
- uvicorn geoguesser_env.server.app:app
194
- ```
195
-
196
- Replaying all 3,037 recorded eval episodes through both settings: episodes
197
- scoring exactly zero fall from **10-46% to 0%** for every model, and the
198
- leaderboard order only changes within the tiers already documented as inside
199
- noise at n=200.
200
-
201
- The terminal observation carries full provenance either way — once the truth is
202
- revealed it can no longer be used to shortcut the episode — so recorded traces
203
- stay complete under `hide_task_identity`.
204
-
205
- ## Baselines
206
-
207
- Measured with `examples/geoguesser_llm_rollout.py` on the committed index,
208
- tasks 0/7/14/21/28, so the numbers are reproducible rather than illustrative.
209
- Five episodes is far too few for a leaderboard; they are a smoke test that the
210
- task is solvable and the reward is discriminative.
211
-
212
- | Model | Mode | Mean reward | Median distance | Within 200 km | Parsed |
213
- |-------|------|------------:|----------------:|--------------:|-------:|
214
- | `claude-sonnet-5` | single-shot | **0.896** | 98 km | 4/5 | 5/5 |
215
- | `claude-sonnet-5` | agentic, tasks 0-5 | 0.539-0.653 | 574-660 km | 3/6 | 6/6 |
216
- | `Qwen/Qwen3.5-9B` | agentic, tasks 0-5 | 0.355 | 1,136 km | 0/5 | 5/6 |
217
- | `Qwen/Qwen3.5-9B:together` | single-shot | 0.277 | 1,139 km | 1/3 | 3/5 |
218
- | `Qwen/Qwen3.5-9B:together` | agentic, 6 turns | 0.304 | 579 km | 0/1 | 1/2 |
219
- | `Qwen/Qwen3.5-9B:together` | agentic, 8k tokens | 0.087 | 2,423 km | 0/1 | **4/4 turns** |
220
- | `Qwen/Qwen3.5-9B:together` | single-shot, 8 eps, 4 parallel | 0.412 | 787 km | 1/6 | 6/8 |
221
-
222
- Sonnet placed two guesses within 2 km. The agentic score sits slightly below
223
- single-shot on the same tasks because looking around costs reward and the extra
224
- views did not always pay for themselves — which is the trade-off the
225
- environment is meant to expose, not a defect.
226
-
227
- Qwen does follow the multi-turn protocol: across the agentic runs it produced
228
- `look`, `move`, `zoom`, `pin` and `guess` actions and navigated up to 68 m down
229
- a road. Two things had to be right first, and both are prompting or plumbing
230
- rather than capability:
231
-
232
- - **Token budget.** Reasoning models put their chain of thought in a separate
233
- `reasoning_content` field and can exhaust the budget before emitting any
234
- content, which looks exactly like a model that cannot see images. At 1,024
235
- tokens Qwen scored 0/5 with empty replies; at 3,500 it followed the protocol
236
- intermittently, failing turns whose reply came back as pure reasoning; at
237
- 8,000 it parsed 4/4 turns. The example defaults to 3,000 and takes
238
- `--max-tokens`.
239
- - **A deadline.** Left to itself Qwen explored until the turn budget ran out and
240
- scored zero. The agentic prompt now warns it explicitly when two turns remain,
241
- after which it committed.
242
-
243
- What remains is accuracy, not plumbing: its guesses landed 452 km, 579 km and
244
- 2,423 km out against Sonnet's 98 km median. It is also 10x slower — 110-193 s
245
- per agentic episode against Sonnet's 10-16 s.
246
-
247
- ```bash
248
- python examples/geoguesser_llm_rollout.py --provider anthropic \
249
- --model claude-sonnet-5 --episodes 5
250
- python examples/geoguesser_llm_rollout.py --provider hf \
251
- --model "Qwen/Qwen3.5-9B:together" --episodes 5 --max-tokens 4000
252
- python examples/geoguesser_llm_rollout.py --provider anthropic \
253
- --mode agentic --episodes 3 --verbose
254
- ```
255
-
256
- ## Parallel rollouts
257
-
258
- An episode is stateful, so concurrent rollouts each need their own environment
259
- instance over the shared read-only index and cache. `--concurrency` does that:
260
-
261
- ```bash
262
- python examples/geoguesser_llm_rollout.py --provider hf \
263
- --model "Qwen/Qwen3.5-9B:together" --episodes 8 --concurrency 4
264
- ```
265
-
266
- Eight Qwen episodes took 113.6 s wall against 334.0 s of summed latency — a
267
- **2.94x speedup** on 4 workers, the shortfall being the provider's own queuing
268
- rather than the environment, which spends ~28 ms on a reset.
269
-
270
- ## Scaling: what the environment can actually supply
271
-
272
- Measured on an 18-core machine with a warm cache and street detail off, one
273
- environment per worker, with a correctness assertion in the loop so an
274
- interference bug cannot masquerade as throughput.
275
-
276
- Per-step cost is dominated by map rendering, not imagery:
277
-
278
- | Step | Cost |
279
- |------|-----:|
280
- | reset, or `look` at 90 deg fov | **29 ms** |
281
- | `look` at 30 deg fov, from the original | 74 ms |
282
- | `place_pin`, a two-panel map | **207 ms** |
283
- | `submit_guess` with the reveal map | 278 ms |
284
-
285
- That shapes the two configurations:
286
-
287
- | Config | Workers | Episodes/s | Notes |
288
- |--------|--------:|-----------:|-------|
289
- | training — views only, no pins, no reveal | 8 threads | **31.7** | 158 env steps/s |
290
- | eval — pins and reveal map | 4 processes | **3.5** | matplotlib is GIL-bound |
291
- | eval — pins and reveal map | 8 threads | 1.9 | threads do not help here |
292
-
293
- Two facts fall out of that. `look` releases the GIL — the reprojection is numpy
294
- and the encode is Pillow — so **threads scale well for view-only work** and
295
- processes only add startup cost. Map rendering is pure Python, so it is
296
- **GIL-bound and needs processes**, which buy about 1.8x before contention.
297
-
298
- `GEOGUESSER_REVEAL_MAP=0` is the single biggest throughput lever: a
299
- training-shaped episode drops from **389 ms to 117 ms**, since the reveal map is
300
- 280 ms that a training run never reads — the reward and the distance are in the
301
- observation either way. Keep it on for evals, traces and the UI.
302
-
303
- Memory is small: about 215 MB for the imports, 100 MB more once the geodata
304
- caches fill, and **~13 MB per additional environment in the same process**. The
305
- Natural Earth layers are process-shared through an LRU cache, so threads are far
306
- cheaper than processes here too.
307
-
308
- ### Rate limits that actually bind
309
-
310
- | API | When it is called | Limit |
311
- |-----|-------------------|-------|
312
- | Mapillary | only on a cache miss | 60,000/min entity, 10,000/min search, 50,000/day tiles |
313
- | Overpass | only a pin below 0.35 deg span | ~10,000/day, 2 concurrent slots per IP |
314
- | the model | every turn | the real constraint |
315
-
316
- With a warm cache the environment makes **no network calls at all**. Default
317
- pins use a 7 degree span, which is above the street threshold, so they do not
318
- touch Overpass either — only a deliberately zoomed pin does. For an eval sweep
319
- set `GEOGUESSER_STREET_DETAIL=0` unless you have pre-warmed, since 100 zoomed
320
- pins against 2 concurrent slots would throttle immediately.
321
-
322
- ### 100-task eval
323
-
324
- Sonnet agentic measured 13.1 s per episode, so 100 episodes cost about 1,310
325
- model-seconds and the concurrency you can use is set by the provider, not by
326
- this environment:
327
-
328
- - concurrency 4 → **~5.5 minutes**
329
- - concurrency 8 → **~2.7 minutes**, if your token-per-minute ceiling allows it
330
-
331
- An agentic episode sends roughly five 640x640 views, about 550 tokens each, plus
332
- a growing text prompt. At a 400k input-tokens-per-minute ceiling that is roughly
333
- 6 episodes in flight before tokens, not latency, become the limit — which is why
334
- 4 to 6 is the practical range for Sonnet. Qwen through the router reached 3.97x
335
- on 6 workers, so 6 to 8 there.
336
-
337
- Meanwhile the environment can supply 3.5 episodes/s in eval configuration
338
- against the 0.3 to 0.6 episodes/s those concurrencies actually consume, so it
339
- has roughly ten times the headroom it needs.
340
-
341
- ### 1000 training steps
342
-
343
- For GRPO with 8 prompts and a group of 16, that is 128 episodes per step and
344
- **128,000 episodes**, or about 640,000 env steps:
345
-
346
- - environment time: 128,000 / 31.7 ≈ **1.1 hours total**, spread across workers
347
- - imagery: nothing, with a warm cache
348
- - storage: 100 tasks warm is 60 MB; 5,000 tasks would be about 1.5 GB of start
349
- frames
350
-
351
- So the environment is not the constraint — 640,000 model calls are, which needs
352
- batched local inference rather than an API. The constraint that *is* ours is
353
- **task diversity**: 128,000 episodes over 100 tasks means each location is seen
354
- 1,280 times, which is memorisation territory. Before a run that long, either
355
- harvest more tasks or add seeded heading augmentation, which multiplies
356
- effective tasks 8 to 12 times from imagery already on disk.
357
-
358
- ## Readiness audit
359
-
360
- `scripts/readiness_check.py` checks the properties that only appear at real
361
- scale and concurrency, rather than on the four committed fixtures:
362
-
363
- ```bash
364
- python scripts/readiness_check.py --full
365
- ```
366
-
367
- ```
368
- [PASS] index integrity 100 tasks, 47 countries, 100 unique sequences
369
- [PASS] all tasks render 100/100 rendered, median reset 28 ms
370
- [PASS] cross-process determinism two subprocesses and this process agree
371
- [PASS] parallel isolation 8 concurrent episodes, each its own task
372
- [PASS] offline with warm cache 12/12 served with fetching disabled
373
- [PASS] reward is discriminative uniform-random 0.029, fixed-point 0.111
374
- [PASS] step latency look 29 ms, pin + map 249 ms
375
- [PASS] one guess per episode a second guess returns reward=None
376
- [PASS] pin never leaks the target 60 pins across 20 tasks revealed nothing
377
- ```
378
-
379
- The reward check matters most: a uniform-random guesser scores **0.029** and the
380
- best trivial constant guess **0.111**, against Sonnet's 0.896. The signal is
381
- measuring geolocation rather than rewarding noise.
382
-
383
- ## Tracing a rollout
384
-
385
- `--trace-dir` records every turn: the image the model saw, what it said, the
386
- action it chose, the environment's reply, steps left and running cost. A second
387
- script renders that as one self-contained HTML page, which is the difference
388
- between knowing the reward and seeing why:
389
-
390
- ```bash
391
- python examples/geoguesser_llm_rollout.py --provider anthropic \
392
- --mode agentic --episodes 6 --trace-dir rollouts
393
- python scripts/render_trace.py rollouts/anthropic_agentic
394
- ```
395
-
396
- Images are written beside the trace rather than inlined, since six agentic
397
- episodes carry around 35 views and a JSONL with those in it is neither readable
398
- nor loadable.
399
-
400
- Every step carries an image, including the guess: a guess returns a **reveal
401
- map** with the guess, the true location and the line between them. Truth is
402
- drawn only there, after scoring. When the guess is more than 25 degrees out the
403
- second panel frames the true location instead of both points, because squashing
404
- a hemisphere into a panel shows nothing.
405
-
406
- ## On the Hub
407
-
408
- | | |
409
- |---|---|
410
- | Space | [`HuggingEnvs/geoguesser-env`](https://huggingface.co/spaces/HuggingEnvs/geoguesser-env) |
411
- | Space (mirror) | [`AdithyaSK/geoguesser-env`](https://huggingface.co/spaces/AdithyaSK/geoguesser-env) |
412
- | Task splits | [`HuggingEnvs/geoguesser-tasks`](https://huggingface.co/datasets/HuggingEnvs/geoguesser-tasks) |
413
- | Imagery | [`HuggingEnvs/geoguesser-panos`](https://huggingface.co/buckets/HuggingEnvs/geoguesser-panos) (Storage Bucket, public, mounted read-only at `/data`) |
414
-
415
- The Space and the bucket deliberately live in different namespaces: moving the
416
- Space should not mean re-uploading 22 GB, so `deploy_hub.py` takes the owner of
417
- each separately (`GEOGUESSER_HF_ORG` and `GEOGUESSER_HF_BUCKET_OWNER`).
418
-
419
- The same client drives either:
420
-
421
- ```python
422
- env = GeoGuesserEnv(base_url="http://localhost:8000") # local
423
- env = GeoGuesserEnv(base_url="https://huggingenvs-geoguesser-env.hf.space") # Space
424
- ```
425
-
426
- Only the file locations differ — locally the indexes and imagery are in the
427
- repo, on the Space they arrive through the bucket mount. Splits, step budget,
428
- street labels and offline enforcement are identical, and verified so: the same
429
- task returns the same image checksums, reward and distance from both.
430
-
431
- Deploy with `scripts/deploy_hub.py --all`. `openenv push` is deliberately not
432
- used: it cannot attach a bucket volume, and its default excludes would upload
433
- 22 GB of panoramas into git.
434
-
435
- ### Local and Space parity
436
-
437
- Verified across a local server and both Spaces, on both splits, driven by the
438
- same client. Every field below is identical in all three:
439
-
440
- | | eval index 42 | train index 1234 |
441
- |---|---|---|
442
- | task id | `eval-00042` | `train-01234` |
443
- | reset image sha256 | `a1447fd8` | `5494a500` |
444
- | look image sha256 | `1c95af16` | `0850e712` |
445
- | move distance | 27.778734090271577 m | 26.196924979479498 m |
446
- | reward, distance | 0.0, 7463.03 km | 0.0, 12647.18 km |
447
-
448
- The guess map is **pixel-identical** too: mean absolute difference 0.000/255
449
- over 1020x390. Only the PNG encoding differs, not the content.
450
-
451
- The one fragile part is the street layer, which comes from Overpass. Overpass
452
- answers a laptop in ~2 s but intermittently returns **504 Gateway Timeout** to
453
- datacenter egress, which once left the Space rendering coarse
454
- Natural-Earth-only maps while a laptop drew the full labelled grid. A single
455
- retry recovers it, and the environment now *reports* the state rather than
456
- degrading silently: observation metadata carries `street_detail` as `on`, `off`
457
- or **`unavailable`**, so a poorer map is visible instead of looking like a
458
- styling choice. Scoring is unaffected either way -- reward is distance-based.
459
- `GEOGUESSER_OSM_CACHE` points the street cache at a mounted, pre-warmed
460
- directory for deployments that cannot reach Overpass at all.
461
-
462
- ## Task API
463
-
464
- The environment implements the core `TaskProvider` protocol, so the Task API
465
- routes core already registers become live. Task discovery is metadata only; it
466
- never starts an episode.
467
-
468
- ```bash
469
- curl localhost:8000/geoguesser_env/splits
470
- # [{"name":"train","type":"train","num_tasks":3452,"default":true},
471
- # {"name":"eval","type":"test","num_tasks":200,"default":false}]
472
-
473
- curl -X POST localhost:8000/geoguesser_env/num_tasks -d '{"split":"eval"}'
474
- curl -X POST localhost:8000/geoguesser_env/task -d '{"split":"eval","index":12}'
475
- ```
476
-
477
- ```python
478
- env.list_splits() # [{"name": "eval", "type": "test", ...}, ...]
479
- env.num_tasks("eval") # 200
480
- env.get_task("eval", 12) # metadata, no coordinates and no country
481
- ```
482
-
483
- **Task specs are deliberately truth-free.** They carry `task_index`, `task_id`,
484
- `split`, `n_frames`, `provider`, `sequence_id` and `offline_ready` — never
485
- coordinates and never the country. A spec travels to whatever orchestrates a
486
- run, and a label sitting in a spec can reach a prompt. The true location is
487
- revealed in observation metadata after the guess, which is the one place it
488
- belongs. Per-country eval breakdowns therefore come from finished episodes, not
489
- from `list_tasks`.
490
-
491
- Splits are invisible to the agent. `RESERVED_TOOL_NAMES` blocks a `reset` MCP
492
- tool, so there is no way for a policy to see or choose its own task — the
493
- "agents cannot reset" invariant.
494
-
495
- ## Collecting evals
496
-
497
- `scripts/collect_eval.py` runs models against a split and records everything
498
- about each episode. One JSONL line per episode, ~13 KB.
499
-
500
- ```bash
501
- export ANTHROPIC_API_KEY=...
502
- python scripts/collect_eval.py --provider anthropic --model claude-sonnet-5 \
503
- --split eval --max-turns 12
504
-
505
- # several endpoints in one run, each at its own concurrency
506
- cp models.example.json models.json # anthropic / openai / HF router / vLLM
507
- python scripts/collect_eval.py --models models.json --split eval
508
- ```
509
-
510
- Output lands in `rollouts/<run_id>/`: `episodes.jsonl` plus a `run.json` with
511
- per-model aggregates (mean reward, median distance, within-1/25/200/750 km,
512
- parse rate, forced guesses, tokens, latency). Runs are **resumable** — re-run the
513
- same `--run-id` and it skips episodes already recorded.
514
-
515
- Each turn carries the camera state **before and after** the action, which is
516
- what makes a rollout re-renderable: a pan from 0 to 270 degrees can only be
517
- animated if both ends are known. It also carries the exact prompt, the raw
518
- reply, separated reasoning, finish reason, token counts, per-attempt errors and
519
- latency, and the provider's own response id.
520
-
521
- **Pixels are not stored.** The environment is deterministic and the panoramas
522
- are local, so a renderer replays the state trajectory instead. Storing them
523
- would cost roughly 14 GB for 200 tasks across six models, all re-derivable. What
524
- *is* stored is a sha256 per observation, so a replay can be *verified* rather
525
- than assumed:
526
-
527
- ```bash
528
- python scripts/verify_replay.py ../../rollouts/<run_id>/episodes.jsonl
529
- # 3 episodes · 19 view turns checked · 19 match · 0 mismatch
530
- # every view turn reproduces byte-for-byte
531
- ```
532
-
533
- That check is not decorative: it exits non-zero on drift, because a video built
534
- from a mismatched trace looks authoritative and shows something the model never
535
- saw. Guess maps are excluded from it by construction — they depend on the
536
- Overpass response of the moment. Use `--save-frames` when you want the literal
537
- bytes anyway.
538
-
539
- ## Rollout video
540
-
541
- Two steps, split where the work naturally divides. Python owns
542
- pixels-from-panoramas, because the gnomonic reprojection already lives here and
543
- is verified byte-exact against the trace. React owns layout, typography and
544
- transitions, because that is where iterating on them is pleasant.
545
-
546
- ```bash
547
- # 1. Frames + timeline.json, straight into the Remotion project's public/
548
- python scripts/render_rollout.py ../../rollouts/<run>/episodes.jsonl --episode 0
549
-
550
- # 2. Compose
551
- cd video && npm install
552
- npx remotion render Rollout out.mp4 --props=public/rollouts/<slug>/timeline.json
553
- npx remotion studio # iterate on the composition live
554
- ```
555
-
556
- **The pan is a real pan.** A `look` from 0 to 270 degrees is not a cut between
557
- two stills: it renders one intermediate gnomonic reprojection per video frame
558
- along the *shortest angular path*, eased like a camera rather than swept
559
- linearly — 350 to 10 degrees pans +20, not -340. Zoom interpolates the field of
560
- view the same way. Verified: a `look` segment produces 29 distinct images, a
561
- `zoom` 24, with no duplicates.
562
-
563
- Layout is a large panorama viewport with a HUD of the state the agent is acting
564
- on (heading, fov, actions left, cost), and a trace pane that reveals turns as
565
- they happen — the active turn carries the model's raw reply, past turns recede.
566
- Then a score card with the guess against the truth.
567
-
568
- Fidelity is checked before anything is composed: every view keyframe is
569
- re-rendered at the size the model saw and compared to the trace's sha256, and a
570
- mismatch aborts. A video that looks authoritative while showing something the
571
- model never saw is worse than no video.
572
-
573
- ## Reproducibility
574
-
575
- ```python
576
- env.reset(split="eval", index=7) # exact task, byte-identical -> GRPO, eval
577
- env.reset(split="train", seed=42) # tasks[42 % n_tasks] -> replay
578
- env.reset() # random task in the default split, split and
579
- # index both recorded in metadata -> UI
580
- ```
581
-
582
- `task_index=` still works as an alias for `index=`, so trajectories recorded
583
- before splits existed still replay. The split is recorded in observation
584
- metadata: without it a bare index is ambiguous across three indexes, and a
585
- trajectory stops being replayable.
586
-
587
- Byte-identical repeats hold because panorama bytes come from a local cache
588
- rather than an expiring CDN URL, reprojection is pure numpy with integer
589
- sampling, and the initial heading is pinned to each panorama's own
590
- `compass_angle`.
591
-
592
- An eval score is only meaningful alongside its provenance — the env version,
593
- the task index, and `GEODATA_VERSION` from `server/render/minimap.py`, since
594
- the bundled vectors determine the reverse-geocode text the agent sees.
595
-
596
- ## Training and collection
597
-
598
- The environment plugs into `openenv.core.harness`, so a rollout function and a
599
- collector come for free:
600
-
601
- ```python
602
- from geoguesser_env import GeoGuesserEnv
603
- from geoguesser_env.harness import GeoGuesserSessionFactory, load_tasks
604
-
605
- tasks = load_tasks("tasks/train_pano_v3.jsonl", repeat=16, split="train")
606
- factory = GeoGuesserSessionFactory(
607
- lambda: GeoGuesserEnv(base_url="http://localhost:8000")
608
- )
609
- ```
610
-
611
- See `examples/geoguesser_rollout.py` for a scripted rollout and
612
- `examples/geoguesser_collect.py` for JSONL collection with resume.
613
-
614
- ## Human play
615
-
616
- A five-round game, 5,000 points a round on the same curve the environment
617
- rewards, so a human score is directly comparable to GeoGuessr intuition and to
618
- the agent's reward (both are shown).
619
-
620
- **A round is one episode with one guess.** The five-round game is a UI wrapper
621
- around five separate episodes; the environment itself never accepts more than
622
- one guess, because `submit_guess` is terminal.
623
-
624
- The page plays through the environment rather than simulating it. It opens the
625
- same WebSocket session API a client uses, calls `reset(task_index=...)`, and
626
- sends every pin, look, zoom and move as a real charged step — so the step
627
- counter, the accumulated cost and the final reward are the environment's own
628
- numbers, not the browser's. A side panel shows the observation stream an agent
629
- would receive, including the environment's own rendered map and views.
630
-
631
- Note that plain REST `/step` builds a fresh environment per request, so a
632
- stateful episode has to run over `/ws`; the Python client does this already.
633
-
634
- ```bash
635
- uv run --project . server
636
- # then open http://localhost:8000/geoguesser/play
637
- ```
638
-
639
- The page stands alone at `/geoguesser/play` and is also embedded in the Gradio
640
- playground's **Custom** tab when the web interface is enabled:
641
-
642
- ```bash
643
- ENABLE_WEB_INTERFACE=true uv run --project . server # http://localhost:8000/web/
644
- ```
645
-
646
- Pick a split and an episode with the `reset(split=)` and `reset(index=)`
647
- controls above the game and press
648
- **load episode**, or **random episode** — the same call an eval harness makes,
649
- so you can replay exactly the episode an agent saw. Those controls live on the
650
- Gradio side because choosing a task is orchestration, not something the player
651
- does mid-round; the page itself reads `?task=` from its URL, so
652
- `/geoguesser/play?task=42` opens that episode directly.
653
-
654
- Drag to look around and scroll to zoom (free, for orientation). The `look()`
655
- and `zoom(30)` buttons run charged environment steps and show what the agent
656
- sees. Arrows, or the arrow keys, walk the road — the main view follows, keeping
657
- your heading. **M** toggles a larger map, **T** the trace panel, **Enter**
658
- submits and then advances. On submit the map takes the screen and draws the
659
- line between guess and truth, exactly like the game; the result bar shows
660
- distance, points, env reward and the true location, and a scoreboard breaks
661
- down all five rounds at the end.
662
-
663
- Panoramas are rendered by Pannellum and the map by MapLibre over OpenFreeMap
664
- tiles — no API key, no request limits. The imagery credit line names the
665
- Mapillary contributor, which the CC-BY-SA licence requires.
666
-
667
- The page has to be a standalone document rather than a Gradio `gr.HTML`
668
- fragment: `gr.HTML` inserts markup without executing `<script>` tags, so the
669
- viewers never initialise and the panel renders blank with no error anywhere.
670
-
671
- Extra routes, all local:
672
-
673
- | Route | Returns |
674
- |-------|---------|
675
- | `/geoguesser/play` | the play page |
676
- | `/geoguesser/tasks` | `{"n_tasks": N}` |
677
- | `/geoguesser/task/{i}` | task metadata, including ground truth for the human UI |
678
- | `/geoguesser/pano/{i}` | the starting equirectangular panorama |
679
- | `/geoguesser/pano/{i}/{frame}` | one frame's panorama, so the viewer follows `move()` |
680
-
681
- The human map uses live tiles; the agent's map stays the offline Natural Earth
682
- render, so the agent keeps a determinism the browser does not need. Note that
683
- `/geoguesser/task/{i}` exposes ground truth — it exists for a person playing in
684
- their own browser, and agent observations still withhold it until the guess.
685
-
686
- ## Configuration
687
-
688
- | Variable | Default | Meaning |
689
- |----------|---------|---------|
690
- | `GEOGUESSER_TASKS_EVAL` | `tasks/eval_pano_v3.jsonl` | Frozen eval split |
691
- | `GEOGUESSER_TASKS_TRAIN` | `tasks/train_pano_v3.jsonl` | Training split |
692
- | `GEOGUESSER_DEFAULT_SPLIT` | `train` | Split `reset()` uses when none is named |
693
- | `GEOGUESSER_INDEX` | `tasks/pano_v1.jsonl` | Legacy single index, used only when no split resolves |
694
- | `GEOGUESSER_CACHE` | `data/panos` | Panorama cache directory |
695
- | `GEOGUESSER_EPISODE_MODE` | `agentic` | `agentic`, `single_shot` or `nmpz` |
696
- | `GEOGUESSER_MAX_STEPS` | `24` | Actions before the episode is cut off |
697
- | `GEOGUESSER_REWARD_MODE` | `coords` | `coords` or `country_only` |
698
- | `GEOGUESSER_HIERARCHICAL` | `0` | Add country and region partial credit |
699
- | `GEOGUESSER_VIEW_SIZE` | `640` | Edge length of rendered views |
700
- | `GEOGUESSER_ALLOW_FETCH` | `1` | Whether a cache miss may reach the API |
701
- | `GEOGUESSER_HIRES_ZOOM` | `1` | Render views at or below 45 deg fov from the original |
702
- | `GEOGUESSER_STREET_DETAIL` | `1` | Fetch real OSM streets below 0.35 deg. Governs Overpass only, independent of `ALLOW_FETCH`, and caches to local disk |
703
- | `GEOGUESSER_REVEAL_MAP` | `1` | Draw the guess-versus-truth map; `0` is 3x faster for training |
704
- | `GEOGUESSER_OSM_CACHE` | `data/geo/osm_cache` | Street-window cache; point at a pre-warmed mount where Overpass is unreachable |
705
- | `MAPILLARY_API_KEY` | — | Needed by the builder, and only on a cache miss |
706
-
707
- ## Data
708
-
709
- Three splits, carved from one 3,673-task pool so contamination is enforced
710
- exactly once, at split time, rather than reasoned about across two harvests:
711
-
712
- | Split | Type | Tasks | Countries | Offline |
713
- |---|---|---|---|---|
714
- | `eval` | `test` | 200 | 73, capped at 4 each | all 24 frames mirrored |
715
- | `train` | `train` | 3,452 | 132 | all 24 frames mirrored |
716
- | `random` | `validation` | 1.2M pool rows | global | no, fetches on demand |
717
-
718
- Separation follows the OSV-5M rule: no shared `sequence_id`, and no training
719
- task within 1 km of an eval task. Frames sit ~3.3 m apart, so holding out an
720
- image while keeping its neighbour holds out nothing. The split script verifies
721
- its own work and exits non-zero if either rule is violated — the committed
722
- split reports 0 shared sequences and a closest train task 1.07 km away.
723
-
724
- | | |
725
- |---|---|
726
- | Frames per task | 23.2 mean (8 min, 24 max), ~3.3 m apart |
727
- | Eval index | 1.2 MB, committed |
728
- | Train index | 20 MB, in the Storage Bucket |
729
- | Imagery | 22 GB for 86k frames, 0.26 MB mean per frame |
730
-
731
- `eval` is committed because a frozen benchmark belongs in version control,
732
- where a change to it shows up in review. The training index and the imagery
733
- live in a Storage Bucket, mounted read-only at `/data` on a Space.
734
-
735
- The `random` split is **not yet implemented** — the plumbing takes arbitrary
736
- named splits, but the pool-backed sampler is still to come.
737
-
738
- Each index is self-contained: every frame's coordinates, heading and capture
739
- date live in the JSONL, so the movement graph resolves offline.
740
- Only image bytes are fetched, and only on a cache miss, because Mapillary
741
- `thumb_*_url` values are expiring signed URLs that cannot be stored.
742
-
743
- Coverage is uneven and worth knowing about. Probing 45 Street-View
744
- coordinates found any Mapillary imagery at 21 and a 360-degree panorama at
745
- only 7, heavily clustered. Panorama-first discovery is therefore the only
746
- approach that works — roughly 5% of probe points yield a usable sequence, so
747
- reaching 100 tasks took two passes with different seeds, merged by
748
- `scripts/merge_task_indexes.py`. Africa and Oceania are thin because 360-degree
749
- contributors are; that is a property of the source, documented rather than
750
- papered over.
751
-
752
- ## Known gaps versus the real game
753
-
754
- Movement follows captured sequences and stops where one ends. There is no
755
- multi-round cumulative score, no wall-clock timer (a step budget stands in for
756
- it), and no satellite layer on the guess map. Coverage hints and web search are
757
- deliberately excluded: the first is a crutch, the second turns the task into
758
- retrieval.
759
-
760
- See [DESIGN.md](DESIGN.md) for the reasoning behind these choices.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/openenv_geoguesser_env.egg-info/SOURCES.txt DELETED
@@ -1,29 +0,0 @@
1
- README.md
2
- __init__.py
3
- client.py
4
- harness.py
5
- models.py
6
- pyproject.toml
7
- ./__init__.py
8
- ./client.py
9
- ./harness.py
10
- ./models.py
11
- openenv_geoguesser_env.egg-info/PKG-INFO
12
- openenv_geoguesser_env.egg-info/SOURCES.txt
13
- openenv_geoguesser_env.egg-info/dependency_links.txt
14
- openenv_geoguesser_env.egg-info/entry_points.txt
15
- openenv_geoguesser_env.egg-info/requires.txt
16
- openenv_geoguesser_env.egg-info/top_level.txt
17
- server/__init__.py
18
- server/app.py
19
- server/geoguesser_environment.py
20
- server/gradio_ui.py
21
- server/parser.py
22
- server/scoring.py
23
- server/backends/__init__.py
24
- server/backends/base.py
25
- server/backends/panorama.py
26
- server/render/__init__.py
27
- server/render/minimap.py
28
- server/render/pano.py
29
- tests/test_geoguesser_env.py
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/openenv_geoguesser_env.egg-info/dependency_links.txt DELETED
@@ -1 +0,0 @@
1
-
 
 
geoguesser_env/openenv_geoguesser_env.egg-info/entry_points.txt DELETED
@@ -1,2 +0,0 @@
1
- [console_scripts]
2
- server = geoguesser_env.server.app:main
 
 
 
geoguesser_env/openenv_geoguesser_env.egg-info/requires.txt DELETED
@@ -1,14 +0,0 @@
1
- openenv>=0.3.1
2
- fastapi>=0.115.0
3
- pydantic>=2.0.0
4
- uvicorn>=0.24.0
5
- fastmcp>=2.0.0
6
- pillow>=10.0.0
7
- numpy>=1.24.0
8
- matplotlib>=3.7.0
9
-
10
- [dev]
11
- pytest>=8.0.0
12
-
13
- [ui]
14
- gradio>=4.0.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
geoguesser_env/openenv_geoguesser_env.egg-info/top_level.txt DELETED
@@ -1 +0,0 @@
1
- geoguesser_env