Spaces:
Running
Running
Remove committed setuptools build/ and egg-info; they shadow the real sources when the Space is pip-installed
Browse files- env/openenv_geoguesser_env.egg-info/PKG-INFO +0 -233
- env/openenv_geoguesser_env.egg-info/SOURCES.txt +0 -21
- env/openenv_geoguesser_env.egg-info/dependency_links.txt +0 -1
- env/openenv_geoguesser_env.egg-info/entry_points.txt +0 -2
- env/openenv_geoguesser_env.egg-info/requires.txt +0 -14
- env/openenv_geoguesser_env.egg-info/top_level.txt +0 -1
- geoguesser_env/build/lib/geoguesser_env/__init__.py +0 -49
- geoguesser_env/build/lib/geoguesser_env/client.py +0 -143
- geoguesser_env/build/lib/geoguesser_env/harness.py +0 -440
- geoguesser_env/build/lib/geoguesser_env/models.py +0 -462
- geoguesser_env/build/lib/geoguesser_env/server/__init__.py +0 -3
- geoguesser_env/build/lib/geoguesser_env/server/app.py +0 -349
- geoguesser_env/build/lib/geoguesser_env/server/backends/__init__.py +0 -7
- geoguesser_env/build/lib/geoguesser_env/server/backends/base.py +0 -127
- geoguesser_env/build/lib/geoguesser_env/server/backends/panorama.py +0 -399
- geoguesser_env/build/lib/geoguesser_env/server/geoguesser_environment.py +0 -998
- geoguesser_env/build/lib/geoguesser_env/server/gradio_ui.py +0 -1044
- geoguesser_env/build/lib/geoguesser_env/server/parser.py +0 -149
- geoguesser_env/build/lib/geoguesser_env/server/render/__init__.py +0 -3
- geoguesser_env/build/lib/geoguesser_env/server/render/minimap.py +0 -824
- geoguesser_env/build/lib/geoguesser_env/server/render/pano.py +0 -108
- geoguesser_env/build/lib/geoguesser_env/server/scoring.py +0 -233
- geoguesser_env/openenv_geoguesser_env.egg-info/PKG-INFO +0 -760
- geoguesser_env/openenv_geoguesser_env.egg-info/SOURCES.txt +0 -29
- geoguesser_env/openenv_geoguesser_env.egg-info/dependency_links.txt +0 -1
- geoguesser_env/openenv_geoguesser_env.egg-info/entry_points.txt +0 -2
- geoguesser_env/openenv_geoguesser_env.egg-info/requires.txt +0 -14
- geoguesser_env/openenv_geoguesser_env.egg-info/top_level.txt +0 -1
env/openenv_geoguesser_env.egg-info/PKG-INFO
DELETED
|
@@ -1,233 +0,0 @@
|
|
| 1 |
-
Metadata-Version: 2.4
|
| 2 |
-
Name: openenv-geoguesser-env
|
| 3 |
-
Version: 0.1.0
|
| 4 |
-
Summary: GeoGuessr-style visual geolocation environment for OpenEnv
|
| 5 |
-
Requires-Python: >=3.10
|
| 6 |
-
Description-Content-Type: text/markdown
|
| 7 |
-
Requires-Dist: openenv>=0.3.1
|
| 8 |
-
Requires-Dist: fastapi>=0.115.0
|
| 9 |
-
Requires-Dist: pydantic>=2.0.0
|
| 10 |
-
Requires-Dist: uvicorn>=0.24.0
|
| 11 |
-
Requires-Dist: fastmcp>=2.0.0
|
| 12 |
-
Requires-Dist: pillow>=10.0.0
|
| 13 |
-
Requires-Dist: numpy>=1.24.0
|
| 14 |
-
Requires-Dist: matplotlib>=3.7.0
|
| 15 |
-
Provides-Extra: ui
|
| 16 |
-
Requires-Dist: gradio>=4.0.0; extra == "ui"
|
| 17 |
-
Provides-Extra: dev
|
| 18 |
-
Requires-Dist: pytest>=8.0.0; extra == "dev"
|
| 19 |
-
|
| 20 |
-
# GeoGuesser
|
| 21 |
-
|
| 22 |
-
A GeoGuessr-style visual geolocation environment. The agent is dropped at an
|
| 23 |
-
unknown street-level location, looks around, walks along the road, pins
|
| 24 |
-
candidate coordinates on a map to check itself, and commits to a final guess.
|
| 25 |
-
Reward is distance-based, using the game's own scoring curve.
|
| 26 |
-
|
| 27 |
-
Independent open-source project, unaffiliated with GeoGuessr AB. Imagery comes
|
| 28 |
-
from Mapillary contributors under CC-BY-SA-4.0.
|
| 29 |
-
|
| 30 |
-
## Quick start
|
| 31 |
-
|
| 32 |
-
```bash
|
| 33 |
-
# 1. Build a task index (needs a free Mapillary token with READ scope)
|
| 34 |
-
export MAPILLARY_API_KEY="MLY|..."
|
| 35 |
-
cd envs/geoguesser_env
|
| 36 |
-
uv run python scripts/build_pano_tasks.py --tasks 100 --frames 24
|
| 37 |
-
|
| 38 |
-
# 2. Run the server
|
| 39 |
-
uv run --project . server # http://localhost:8000
|
| 40 |
-
```
|
| 41 |
-
|
| 42 |
-
```python
|
| 43 |
-
from geoguesser_env import GeoGuesserEnv, GuessAction, LookAction, PinAction
|
| 44 |
-
|
| 45 |
-
env = GeoGuesserEnv(base_url="http://localhost:8000")
|
| 46 |
-
|
| 47 |
-
result = env.reset(task_index=7) # byte-identical on repeat
|
| 48 |
-
print(result.observation.prompt)
|
| 49 |
-
|
| 50 |
-
result = env.step(LookAction(heading_deg=90, fov_deg=45))
|
| 51 |
-
result = env.step(PinAction(lat=-16.5, lon=-68.1))
|
| 52 |
-
print(result.observation.feedback)
|
| 53 |
-
# Pin 1 placed at -16.5000, -68.1000 - Bolivia (South America).
|
| 54 |
-
# Nearest major city: La Paz, ~5 km E. 10 actions left.
|
| 55 |
-
|
| 56 |
-
result = env.step(GuessAction(response="Altiplano. <guess>-16.49, -68.12</guess>"))
|
| 57 |
-
print(result.reward, result.observation.distance_km)
|
| 58 |
-
```
|
| 59 |
-
|
| 60 |
-
## Tools
|
| 61 |
-
|
| 62 |
-
| Tool | What it does | Cost |
|
| 63 |
-
|------|--------------|------|
|
| 64 |
-
| `look(heading_deg, pitch_deg, fov_deg)` | Render a view. Heading is absolute, `0` is true north | −0.01 |
|
| 65 |
-
| `pan(delta_deg)` | Turn relative to the current heading | −0.01 |
|
| 66 |
-
| `zoom(fov_deg)` | Narrow the field of view; around 30 reads distant signs | −0.01 |
|
| 67 |
-
| `move(direction, meters)` | Walk the captured road; reports distance actually travelled | −0.05 |
|
| 68 |
-
| `place_pin(lat, lon, label)` | Pin a candidate and see where it falls on the map | −0.02 |
|
| 69 |
-
| `view_map(lat, lon, span_deg)` | Pan and zoom the map without pinning | −0.01 |
|
| 70 |
-
| `list_pins()` / `clear_pins()` | Review or drop candidates | free |
|
| 71 |
-
| `measure(lat_a, lon_a, lat_b, lon_b)` | Distance between two of your own points | free |
|
| 72 |
-
| `reverse_geocode(lat, lon)` | Name the country and nearest city at a coordinate | free |
|
| 73 |
-
| `submit_guess(lat, lon, ...)` | Commit the answer. Terminal | — |
|
| 74 |
-
|
| 75 |
-
Tools the backend cannot serve are **not registered**, so the agent never sees
|
| 76 |
-
a tool that always fails.
|
| 77 |
-
|
| 78 |
-
### Pinning tells you where you pointed, not whether you are right
|
| 79 |
-
|
| 80 |
-
`place_pin` returns a rendered map and a description of the pinned location:
|
| 81 |
-
country, subregion, nearest city with distance and bearing, and the distance
|
| 82 |
-
to the agent's own earlier pins. It reveals nothing about the target.
|
| 83 |
-
|
| 84 |
-
That restraint is deliberate. Any signal about the truth — a distance, a
|
| 85 |
-
warmer/colder hint — would make binary search the optimal policy, and the
|
| 86 |
-
environment would measure bisection rather than geographic reasoning. Distance
|
| 87 |
-
and score arrive only from `submit_guess`.
|
| 88 |
-
|
| 89 |
-
## Reward
|
| 90 |
-
|
| 91 |
-
```
|
| 92 |
-
geo = exp(-distance_km / 1492.7) # GeoGuessr's curve, in [0, 1]
|
| 93 |
-
partial = 0.15 * country_hit + 0.10 * region_hit # when hierarchical
|
| 94 |
-
cost = 0.01*looks + 0.01*maps + 0.02*pins + 0.05*moves
|
| 95 |
-
reward = clip(geo + partial, 0, 1) - cost
|
| 96 |
-
```
|
| 97 |
-
|
| 98 |
-
An unparseable or out-of-range guess scores `0.0` and says why. Parsing
|
| 99 |
-
accepts what models actually emit: decimal pairs, DMS (`48°51'29"N`), labelled
|
| 100 |
-
`lat:`/`lon:`, JSON, and `<guess>` tags.
|
| 101 |
-
|
| 102 |
-
## Reproducibility
|
| 103 |
-
|
| 104 |
-
```python
|
| 105 |
-
env.reset(task_index=7) # exact task, byte-identical observation -> GRPO, eval
|
| 106 |
-
env.reset(seed=42) # tasks[42 % n_tasks] -> replay
|
| 107 |
-
env.reset() # random task, index in metadata -> UI
|
| 108 |
-
```
|
| 109 |
-
|
| 110 |
-
Byte-identical repeats hold because panorama bytes come from a local cache
|
| 111 |
-
rather than an expiring CDN URL, reprojection is pure numpy with integer
|
| 112 |
-
sampling, and the initial heading is pinned to each panorama's own
|
| 113 |
-
`compass_angle`.
|
| 114 |
-
|
| 115 |
-
An eval score is only meaningful alongside its provenance — the env version,
|
| 116 |
-
the task index, and `GEODATA_VERSION` from `server/render/minimap.py`, since
|
| 117 |
-
the bundled vectors determine the reverse-geocode text the agent sees.
|
| 118 |
-
|
| 119 |
-
## Training and collection
|
| 120 |
-
|
| 121 |
-
The environment plugs into `openenv.core.harness`, so a rollout function and a
|
| 122 |
-
collector come for free:
|
| 123 |
-
|
| 124 |
-
```python
|
| 125 |
-
from geoguesser_env import GeoGuesserEnv
|
| 126 |
-
from geoguesser_env.harness import GeoGuesserSessionFactory, load_tasks
|
| 127 |
-
|
| 128 |
-
tasks = load_tasks("tasks/pano_v1.jsonl", repeat=16) # GRPO group of 16
|
| 129 |
-
factory = GeoGuesserSessionFactory(
|
| 130 |
-
lambda: GeoGuesserEnv(base_url="http://localhost:8000")
|
| 131 |
-
)
|
| 132 |
-
```
|
| 133 |
-
|
| 134 |
-
See `examples/geoguesser_rollout.py` for a scripted rollout and
|
| 135 |
-
`examples/geoguesser_collect.py` for JSONL collection with resume.
|
| 136 |
-
|
| 137 |
-
## Human play
|
| 138 |
-
|
| 139 |
-
A five-round game, 5,000 points a round on the same curve the environment
|
| 140 |
-
rewards, so a human score is directly comparable to GeoGuessr intuition and to
|
| 141 |
-
the agent's reward (both are shown).
|
| 142 |
-
|
| 143 |
-
```bash
|
| 144 |
-
uv run --project . server
|
| 145 |
-
# then open http://localhost:8000/geoguesser/play
|
| 146 |
-
```
|
| 147 |
-
|
| 148 |
-
The page stands alone at `/geoguesser/play` and is also embedded in the Gradio
|
| 149 |
-
playground's **Custom** tab when the web interface is enabled:
|
| 150 |
-
|
| 151 |
-
```bash
|
| 152 |
-
ENABLE_WEB_INTERFACE=true uv run --project . server # http://localhost:8000/web/
|
| 153 |
-
```
|
| 154 |
-
|
| 155 |
-
Drag to look around, scroll to zoom, **M** toggles a larger map, **Enter**
|
| 156 |
-
submits and then advances. On submit the map takes the screen and draws the
|
| 157 |
-
line between guess and truth, exactly like the game; the result bar shows
|
| 158 |
-
distance, points, env reward and the true location, and a scoreboard breaks
|
| 159 |
-
down all five rounds at the end.
|
| 160 |
-
|
| 161 |
-
Panoramas are rendered by Pannellum and the map by MapLibre over OpenFreeMap
|
| 162 |
-
tiles — no API key, no request limits. The imagery credit line names the
|
| 163 |
-
Mapillary contributor, which the CC-BY-SA licence requires.
|
| 164 |
-
|
| 165 |
-
The page has to be a standalone document rather than a Gradio `gr.HTML`
|
| 166 |
-
fragment: `gr.HTML` inserts markup without executing `<script>` tags, so the
|
| 167 |
-
viewers never initialise and the panel renders blank with no error anywhere.
|
| 168 |
-
|
| 169 |
-
Extra routes, all local:
|
| 170 |
-
|
| 171 |
-
| Route | Returns |
|
| 172 |
-
|-------|---------|
|
| 173 |
-
| `/geoguesser/play` | the play page |
|
| 174 |
-
| `/geoguesser/tasks` | `{"n_tasks": N}` |
|
| 175 |
-
| `/geoguesser/task/{i}` | task metadata, including ground truth for the human UI |
|
| 176 |
-
| `/geoguesser/pano/{i}` | the raw equirectangular panorama |
|
| 177 |
-
|
| 178 |
-
The human map uses live tiles; the agent's map stays the offline Natural Earth
|
| 179 |
-
render, so the agent keeps a determinism the browser does not need. Note that
|
| 180 |
-
`/geoguesser/task/{i}` exposes ground truth — it exists for a person playing in
|
| 181 |
-
their own browser, and agent observations still withhold it until the guess.
|
| 182 |
-
|
| 183 |
-
## Configuration
|
| 184 |
-
|
| 185 |
-
| Variable | Default | Meaning |
|
| 186 |
-
|----------|---------|---------|
|
| 187 |
-
| `GEOGUESSER_INDEX` | `tasks/pano_v1.jsonl` | Task index to load |
|
| 188 |
-
| `GEOGUESSER_CACHE` | `data/panos` | Panorama cache directory |
|
| 189 |
-
| `GEOGUESSER_EPISODE_MODE` | `agentic` | `agentic`, `single_shot` or `nmpz` |
|
| 190 |
-
| `GEOGUESSER_MAX_STEPS` | `12` | Actions before the episode is cut off |
|
| 191 |
-
| `GEOGUESSER_REWARD_MODE` | `coords` | `coords` or `country_only` |
|
| 192 |
-
| `GEOGUESSER_HIERARCHICAL` | `0` | Add country and region partial credit |
|
| 193 |
-
| `GEOGUESSER_VIEW_SIZE` | `640` | Edge length of rendered views |
|
| 194 |
-
| `GEOGUESSER_ALLOW_FETCH` | `1` | Whether a cache miss may reach the API |
|
| 195 |
-
| `MAPILLARY_API_KEY` | — | Needed by the builder, and only on a cache miss |
|
| 196 |
-
|
| 197 |
-
## Data
|
| 198 |
-
|
| 199 |
-
The committed index holds **100 tasks across 47 countries and all six
|
| 200 |
-
continents**, one per Mapillary sequence, captured between 2017 and 2026:
|
| 201 |
-
|
| 202 |
-
| | |
|
| 203 |
-
|---|---|
|
| 204 |
-
| Tasks | 100 (`task_index` 0-99) |
|
| 205 |
-
| Countries | 47, capped at 4 tasks each |
|
| 206 |
-
| Continents | Europe 34, Asia 27, South America 23, North America 11, Oceania 3, Africa 2 |
|
| 207 |
-
| Frames per task | 23.2 mean (3 min, 24 max), ~3.3 m apart |
|
| 208 |
-
| Index size | 417 KB, committed |
|
| 209 |
-
| Cache | 37 MB for the 100 start frames; ~0.6 GB fully warmed |
|
| 210 |
-
|
| 211 |
-
The index is self-contained: every frame's coordinates, heading and capture
|
| 212 |
-
date live in `tasks/pano_v1.jsonl`, so the movement graph resolves offline.
|
| 213 |
-
Only image bytes are fetched, and only on a cache miss, because Mapillary
|
| 214 |
-
`thumb_*_url` values are expiring signed URLs that cannot be stored.
|
| 215 |
-
|
| 216 |
-
Coverage is uneven and worth knowing about. Probing 45 Street-View
|
| 217 |
-
coordinates found any Mapillary imagery at 21 and a 360-degree panorama at
|
| 218 |
-
only 7, heavily clustered. Panorama-first discovery is therefore the only
|
| 219 |
-
approach that works — roughly 5% of probe points yield a usable sequence, so
|
| 220 |
-
reaching 100 tasks took two passes with different seeds, merged by
|
| 221 |
-
`scripts/merge_task_indexes.py`. Africa and Oceania are thin because 360-degree
|
| 222 |
-
contributors are; that is a property of the source, documented rather than
|
| 223 |
-
papered over.
|
| 224 |
-
|
| 225 |
-
## Known gaps versus the real game
|
| 226 |
-
|
| 227 |
-
Movement follows captured sequences and stops where one ends. There is no
|
| 228 |
-
multi-round cumulative score, no wall-clock timer (a step budget stands in for
|
| 229 |
-
it), and no satellite layer on the guess map. Coverage hints and web search are
|
| 230 |
-
deliberately excluded: the first is a crutch, the second turns the task into
|
| 231 |
-
retrieval.
|
| 232 |
-
|
| 233 |
-
See [DESIGN.md](DESIGN.md) for the reasoning behind these choices.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
env/openenv_geoguesser_env.egg-info/SOURCES.txt
DELETED
|
@@ -1,21 +0,0 @@
|
|
| 1 |
-
README.md
|
| 2 |
-
pyproject.toml
|
| 3 |
-
openenv_geoguesser_env.egg-info/PKG-INFO
|
| 4 |
-
openenv_geoguesser_env.egg-info/SOURCES.txt
|
| 5 |
-
openenv_geoguesser_env.egg-info/dependency_links.txt
|
| 6 |
-
openenv_geoguesser_env.egg-info/entry_points.txt
|
| 7 |
-
openenv_geoguesser_env.egg-info/requires.txt
|
| 8 |
-
openenv_geoguesser_env.egg-info/top_level.txt
|
| 9 |
-
server/__init__.py
|
| 10 |
-
server/app.py
|
| 11 |
-
server/geoguesser_environment.py
|
| 12 |
-
server/gradio_ui.py
|
| 13 |
-
server/parser.py
|
| 14 |
-
server/scoring.py
|
| 15 |
-
server/backends/__init__.py
|
| 16 |
-
server/backends/base.py
|
| 17 |
-
server/backends/panorama.py
|
| 18 |
-
server/render/__init__.py
|
| 19 |
-
server/render/minimap.py
|
| 20 |
-
server/render/pano.py
|
| 21 |
-
tests/test_geoguesser_env.py
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
env/openenv_geoguesser_env.egg-info/dependency_links.txt
DELETED
|
@@ -1 +0,0 @@
|
|
| 1 |
-
|
|
|
|
|
|
env/openenv_geoguesser_env.egg-info/entry_points.txt
DELETED
|
@@ -1,2 +0,0 @@
|
|
| 1 |
-
[console_scripts]
|
| 2 |
-
server = server.app:main
|
|
|
|
|
|
|
|
|
env/openenv_geoguesser_env.egg-info/requires.txt
DELETED
|
@@ -1,14 +0,0 @@
|
|
| 1 |
-
openenv>=0.3.1
|
| 2 |
-
fastapi>=0.115.0
|
| 3 |
-
pydantic>=2.0.0
|
| 4 |
-
uvicorn>=0.24.0
|
| 5 |
-
fastmcp>=2.0.0
|
| 6 |
-
pillow>=10.0.0
|
| 7 |
-
numpy>=1.24.0
|
| 8 |
-
matplotlib>=3.7.0
|
| 9 |
-
|
| 10 |
-
[dev]
|
| 11 |
-
pytest>=8.0.0
|
| 12 |
-
|
| 13 |
-
[ui]
|
| 14 |
-
gradio>=4.0.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
env/openenv_geoguesser_env.egg-info/top_level.txt
DELETED
|
@@ -1 +0,0 @@
|
|
| 1 |
-
server
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/__init__.py
DELETED
|
@@ -1,49 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""GeoGuesser: a GeoGuessr-style visual geolocation environment.
|
| 4 |
-
|
| 5 |
-
Independent open-source project, unaffiliated with GeoGuessr AB. Imagery comes
|
| 6 |
-
from Mapillary contributors under CC-BY-SA-4.0.
|
| 7 |
-
"""
|
| 8 |
-
|
| 9 |
-
from .client import GeoGuesserEnv
|
| 10 |
-
from .models import (
|
| 11 |
-
EpisodeMode,
|
| 12 |
-
from_wire,
|
| 13 |
-
GeoGuesserAction,
|
| 14 |
-
GeoGuesserObservation,
|
| 15 |
-
GeoGuesserState,
|
| 16 |
-
GuessAction,
|
| 17 |
-
LookAction,
|
| 18 |
-
MeasureAction,
|
| 19 |
-
MoveAction,
|
| 20 |
-
PanAction,
|
| 21 |
-
Pin,
|
| 22 |
-
PinAction,
|
| 23 |
-
RewardMode,
|
| 24 |
-
to_wire,
|
| 25 |
-
TypedAction,
|
| 26 |
-
ViewMapAction,
|
| 27 |
-
ZoomAction,
|
| 28 |
-
)
|
| 29 |
-
|
| 30 |
-
__all__ = [
|
| 31 |
-
"EpisodeMode",
|
| 32 |
-
"GeoGuesserAction",
|
| 33 |
-
"GeoGuesserEnv",
|
| 34 |
-
"GeoGuesserObservation",
|
| 35 |
-
"GeoGuesserState",
|
| 36 |
-
"GuessAction",
|
| 37 |
-
"LookAction",
|
| 38 |
-
"MeasureAction",
|
| 39 |
-
"MoveAction",
|
| 40 |
-
"PanAction",
|
| 41 |
-
"Pin",
|
| 42 |
-
"PinAction",
|
| 43 |
-
"RewardMode",
|
| 44 |
-
"TypedAction",
|
| 45 |
-
"ViewMapAction",
|
| 46 |
-
"ZoomAction",
|
| 47 |
-
"from_wire",
|
| 48 |
-
"to_wire",
|
| 49 |
-
]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/client.py
DELETED
|
@@ -1,143 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""HTTP client for the GeoGuesser environment."""
|
| 4 |
-
|
| 5 |
-
from __future__ import annotations
|
| 6 |
-
|
| 7 |
-
from typing import Any, Dict
|
| 8 |
-
|
| 9 |
-
from openenv.core.client_types import StepResult
|
| 10 |
-
from openenv.core.env_client import EnvClient
|
| 11 |
-
|
| 12 |
-
from .models import (
|
| 13 |
-
GeoGuesserAction,
|
| 14 |
-
GeoGuesserObservation,
|
| 15 |
-
GeoGuesserState,
|
| 16 |
-
to_wire,
|
| 17 |
-
TypedAction,
|
| 18 |
-
)
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
ENV_NAME = "geoguesser_env"
|
| 22 |
-
"""Task API routes are namespaced by the `env_name` the server registers."""
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
class GeoGuesserEnv(
|
| 26 |
-
EnvClient[GeoGuesserAction, GeoGuesserObservation, GeoGuesserState]
|
| 27 |
-
):
|
| 28 |
-
"""
|
| 29 |
-
Client for a running GeoGuesser environment server.
|
| 30 |
-
|
| 31 |
-
Examples:
|
| 32 |
-
|
| 33 |
-
```python
|
| 34 |
-
env = GeoGuesserEnv(base_url="http://localhost:8000")
|
| 35 |
-
result = env.reset(task_index=0)
|
| 36 |
-
result = env.step(LookAction(heading_deg=90))
|
| 37 |
-
result = env.step(GuessAction(response("<guess>55.67, 12.57</guess>")))
|
| 38 |
-
print(result.reward, result.observation.distance_km)
|
| 39 |
-
```
|
| 40 |
-
"""
|
| 41 |
-
|
| 42 |
-
def _step_payload(self, action: GeoGuesserAction | TypedAction) -> Dict[str, Any]:
|
| 43 |
-
"""
|
| 44 |
-
Serialise an action for the `/step` endpoint.
|
| 45 |
-
|
| 46 |
-
Typed actions are flattened to the single wire schema the server
|
| 47 |
-
declares; a wire action passes through unchanged.
|
| 48 |
-
"""
|
| 49 |
-
wire = action if isinstance(action, GeoGuesserAction) else to_wire(action)
|
| 50 |
-
return wire.model_dump(exclude_none=True)
|
| 51 |
-
|
| 52 |
-
def _parse_result(self, response: Dict[str, Any]) -> StepResult:
|
| 53 |
-
"""Build a [`StepResult`] from a `/step` or `/reset` response."""
|
| 54 |
-
observation = GeoGuesserObservation(**response["observation"])
|
| 55 |
-
return StepResult(
|
| 56 |
-
observation=observation,
|
| 57 |
-
reward=response.get("reward"),
|
| 58 |
-
done=response.get("done", False),
|
| 59 |
-
)
|
| 60 |
-
|
| 61 |
-
def _parse_state(self, response: Dict[str, Any]) -> GeoGuesserState:
|
| 62 |
-
"""Build a [`GeoGuesserState`] from a `/state` response."""
|
| 63 |
-
return GeoGuesserState(**response)
|
| 64 |
-
|
| 65 |
-
def reset(
|
| 66 |
-
self,
|
| 67 |
-
task_index: int | None = None,
|
| 68 |
-
split: str | None = None,
|
| 69 |
-
index: int | None = None,
|
| 70 |
-
**kwargs: Any,
|
| 71 |
-
) -> StepResult:
|
| 72 |
-
"""
|
| 73 |
-
Start an episode.
|
| 74 |
-
|
| 75 |
-
Args:
|
| 76 |
-
task_index (`int`, *optional*):
|
| 77 |
-
Deprecated alias for `index`, kept so existing callers and
|
| 78 |
-
recorded trajectories keep working.
|
| 79 |
-
split (`str`, *optional*):
|
| 80 |
-
Which split to draw from, as named by [`list_splits`]. Defaults
|
| 81 |
-
to the server's default split.
|
| 82 |
-
index (`int`, *optional*):
|
| 83 |
-
Exact task to play within `split`. Repeated calls with the same
|
| 84 |
-
split and index yield byte-identical observations, which is what
|
| 85 |
-
a GRPO group needs. Omit it, and `seed`, for a random task.
|
| 86 |
-
|
| 87 |
-
Returns:
|
| 88 |
-
[`StepResult`]: The opening observation.
|
| 89 |
-
"""
|
| 90 |
-
if index is None:
|
| 91 |
-
index = task_index
|
| 92 |
-
if split is not None:
|
| 93 |
-
kwargs["split"] = split
|
| 94 |
-
if index is not None:
|
| 95 |
-
kwargs["index"] = index
|
| 96 |
-
return super().reset(**kwargs)
|
| 97 |
-
|
| 98 |
-
def list_splits(self) -> list[dict[str, Any]]:
|
| 99 |
-
"""
|
| 100 |
-
Which splits the server offers, and how many tasks each holds.
|
| 101 |
-
|
| 102 |
-
Core exposes the Task API over HTTP but ships no client for it, so this
|
| 103 |
-
posts to the routes directly.
|
| 104 |
-
|
| 105 |
-
Returns:
|
| 106 |
-
`list[dict]`: Split descriptors with `name`, `type`, `num_tasks` and
|
| 107 |
-
`default`.
|
| 108 |
-
"""
|
| 109 |
-
return self._task_api("splits", method="GET")
|
| 110 |
-
|
| 111 |
-
def num_tasks(self, split: str) -> int:
|
| 112 |
-
"""How many tasks a split holds."""
|
| 113 |
-
return int(self._task_api("num_tasks", {"split": split})["num_tasks"])
|
| 114 |
-
|
| 115 |
-
def get_task(self, split: str, index: int) -> dict[str, Any]:
|
| 116 |
-
"""
|
| 117 |
-
Describe one task without starting an episode.
|
| 118 |
-
|
| 119 |
-
The spec carries no coordinates and no country: it is metadata for
|
| 120 |
-
whatever orchestrates a run, not a label source.
|
| 121 |
-
"""
|
| 122 |
-
return self._task_api("task", {"split": split, "index": index})["task"]
|
| 123 |
-
|
| 124 |
-
def _task_api(
|
| 125 |
-
self,
|
| 126 |
-
route: str,
|
| 127 |
-
payload: dict[str, Any] | None = None,
|
| 128 |
-
method: str = "POST",
|
| 129 |
-
) -> Any:
|
| 130 |
-
"""Call one core Task API route on this environment."""
|
| 131 |
-
import json
|
| 132 |
-
import urllib.request
|
| 133 |
-
|
| 134 |
-
url = f"{self.base_url.rstrip('/')}/{ENV_NAME}/{route}"
|
| 135 |
-
data = None if payload is None else json.dumps(payload).encode()
|
| 136 |
-
request = urllib.request.Request(
|
| 137 |
-
url,
|
| 138 |
-
data=data,
|
| 139 |
-
method=method,
|
| 140 |
-
headers={"Content-Type": "application/json"},
|
| 141 |
-
)
|
| 142 |
-
with urllib.request.urlopen(request, timeout=30) as response:
|
| 143 |
-
return json.loads(response.read())
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/harness.py
DELETED
|
@@ -1,440 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Harness-oriented GeoGuesser session adapters.
|
| 4 |
-
|
| 5 |
-
Follows the pattern in `reasoning_gym_env.harness`: expose a `GeoGuesserEnv`
|
| 6 |
-
client as a `ResourceSession` driven through MCP-style tools, so it plugs into
|
| 7 |
-
`openenv.core.harness` unchanged.
|
| 8 |
-
|
| 9 |
-
That single adapter is what makes the rest work without extra code:
|
| 10 |
-
|
| 11 |
-
- `CollectRunner(tasks=...)` walks a task list into a JSONL rollout dataset,
|
| 12 |
-
with resume, recording which task produced each episode.
|
| 13 |
-
- `build_harness_rollout_func(...)` yields a TRL-compatible rollout function
|
| 14 |
-
where each prompt *is* a task — the GRPO path.
|
| 15 |
-
- `EvalConfig` / `EvalResult` carry the provenance an eval score needs.
|
| 16 |
-
|
| 17 |
-
A task is a plain dict, so it survives serialisation into an `EpisodeRecord`:
|
| 18 |
-
|
| 19 |
-
```python
|
| 20 |
-
{"task_index": 7, "task_id": "mly-0007"}
|
| 21 |
-
```
|
| 22 |
-
"""
|
| 23 |
-
|
| 24 |
-
from __future__ import annotations
|
| 25 |
-
|
| 26 |
-
import json
|
| 27 |
-
import pathlib
|
| 28 |
-
from typing import Any, Callable, Iterator
|
| 29 |
-
|
| 30 |
-
from openenv.core.env_server.mcp_types import Tool
|
| 31 |
-
from openenv.core.harness import (
|
| 32 |
-
ResourceSessionFactory,
|
| 33 |
-
StepEnvSessionAdapter,
|
| 34 |
-
ToolResult,
|
| 35 |
-
VerifyResult,
|
| 36 |
-
)
|
| 37 |
-
|
| 38 |
-
from .client import GeoGuesserEnv
|
| 39 |
-
from .models import (
|
| 40 |
-
GuessAction,
|
| 41 |
-
LookAction,
|
| 42 |
-
MeasureAction,
|
| 43 |
-
MoveAction,
|
| 44 |
-
PanAction,
|
| 45 |
-
PinAction,
|
| 46 |
-
to_wire,
|
| 47 |
-
ViewMapAction,
|
| 48 |
-
ZoomAction,
|
| 49 |
-
)
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
def _number(description: str) -> dict[str, Any]:
|
| 53 |
-
return {"type": "number", "description": description}
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
GEOGUESSER_TOOLS: list[Tool] = [
|
| 57 |
-
Tool(
|
| 58 |
-
name="look",
|
| 59 |
-
description=(
|
| 60 |
-
"Look in a direction from where you stand. heading_deg is "
|
| 61 |
-
"absolute, 0 being true north. Smaller fov_deg zooms in."
|
| 62 |
-
),
|
| 63 |
-
input_schema={
|
| 64 |
-
"type": "object",
|
| 65 |
-
"properties": {
|
| 66 |
-
"heading_deg": _number("Compass heading in degrees."),
|
| 67 |
-
"pitch_deg": _number("Vertical angle; positive looks up."),
|
| 68 |
-
"fov_deg": _number("Field of view in degrees, 10 to 120."),
|
| 69 |
-
},
|
| 70 |
-
"required": ["heading_deg"],
|
| 71 |
-
},
|
| 72 |
-
),
|
| 73 |
-
Tool(
|
| 74 |
-
name="pan",
|
| 75 |
-
description="Turn relative to your current heading; positive turns right.",
|
| 76 |
-
input_schema={
|
| 77 |
-
"type": "object",
|
| 78 |
-
"properties": {"delta_deg": _number("Degrees to turn.")},
|
| 79 |
-
"required": ["delta_deg"],
|
| 80 |
-
},
|
| 81 |
-
),
|
| 82 |
-
Tool(
|
| 83 |
-
name="zoom",
|
| 84 |
-
description="Change field of view without turning. Around 30 reads distant signs.",
|
| 85 |
-
input_schema={
|
| 86 |
-
"type": "object",
|
| 87 |
-
"properties": {"fov_deg": _number("New field of view in degrees.")},
|
| 88 |
-
"required": ["fov_deg"],
|
| 89 |
-
},
|
| 90 |
-
),
|
| 91 |
-
Tool(
|
| 92 |
-
name="move",
|
| 93 |
-
description=(
|
| 94 |
-
"Walk along the captured road. Reports how far you actually "
|
| 95 |
-
"travelled, since frame spacing is irregular."
|
| 96 |
-
),
|
| 97 |
-
input_schema={
|
| 98 |
-
"type": "object",
|
| 99 |
-
"properties": {
|
| 100 |
-
"direction": {
|
| 101 |
-
"type": "string",
|
| 102 |
-
"enum": ["forward", "backward"],
|
| 103 |
-
"description": "Which way to walk.",
|
| 104 |
-
},
|
| 105 |
-
"meters": _number("Requested distance in metres."),
|
| 106 |
-
},
|
| 107 |
-
"required": ["direction"],
|
| 108 |
-
},
|
| 109 |
-
),
|
| 110 |
-
Tool(
|
| 111 |
-
name="place_pin",
|
| 112 |
-
description=(
|
| 113 |
-
"Pin a candidate coordinate and see where it falls on the map. "
|
| 114 |
-
"Tells you what is at that coordinate. It says nothing about "
|
| 115 |
-
"whether you are right."
|
| 116 |
-
),
|
| 117 |
-
input_schema={
|
| 118 |
-
"type": "object",
|
| 119 |
-
"properties": {
|
| 120 |
-
"lat": _number("Latitude of the candidate."),
|
| 121 |
-
"lon": _number("Longitude of the candidate."),
|
| 122 |
-
"label": {"type": "string", "description": "Optional note."},
|
| 123 |
-
"span_deg": _number(
|
| 124 |
-
"Half-width of the returned map in degrees. Below about 4 "
|
| 125 |
-
"the map adds roads, urban areas and town names."
|
| 126 |
-
),
|
| 127 |
-
},
|
| 128 |
-
"required": ["lat", "lon"],
|
| 129 |
-
},
|
| 130 |
-
),
|
| 131 |
-
Tool(
|
| 132 |
-
name="view_map",
|
| 133 |
-
description="Pan and zoom the map without placing a pin.",
|
| 134 |
-
input_schema={
|
| 135 |
-
"type": "object",
|
| 136 |
-
"properties": {
|
| 137 |
-
"lat": _number("Latitude at the centre of the view."),
|
| 138 |
-
"lon": _number("Longitude at the centre of the view."),
|
| 139 |
-
"span_deg": _number("Half-width of the window in degrees."),
|
| 140 |
-
},
|
| 141 |
-
"required": ["lat", "lon"],
|
| 142 |
-
},
|
| 143 |
-
),
|
| 144 |
-
Tool(
|
| 145 |
-
name="measure",
|
| 146 |
-
description="Distance in km between two coordinates of your own choosing. Free.",
|
| 147 |
-
input_schema={
|
| 148 |
-
"type": "object",
|
| 149 |
-
"properties": {
|
| 150 |
-
"lat_a": _number("Latitude of the first point."),
|
| 151 |
-
"lon_a": _number("Longitude of the first point."),
|
| 152 |
-
"lat_b": _number("Latitude of the second point."),
|
| 153 |
-
"lon_b": _number("Longitude of the second point."),
|
| 154 |
-
},
|
| 155 |
-
"required": ["lat_a", "lon_a", "lat_b", "lon_b"],
|
| 156 |
-
},
|
| 157 |
-
),
|
| 158 |
-
Tool(
|
| 159 |
-
name="submit_guess",
|
| 160 |
-
description="Commit your final answer. This ends the episode.",
|
| 161 |
-
input_schema={
|
| 162 |
-
"type": "object",
|
| 163 |
-
"properties": {
|
| 164 |
-
"lat": _number("Latitude of your guess."),
|
| 165 |
-
"lon": _number("Longitude of your guess."),
|
| 166 |
-
"country": {
|
| 167 |
-
"type": "string",
|
| 168 |
-
"description": "Optional ISO-3166 alpha-2 code or country name.",
|
| 169 |
-
},
|
| 170 |
-
"confidence": _number("Optional confidence in [0, 1]."),
|
| 171 |
-
"reasoning": {
|
| 172 |
-
"type": "string",
|
| 173 |
-
"description": "Optional rationale, recorded but not scored.",
|
| 174 |
-
},
|
| 175 |
-
},
|
| 176 |
-
"required": ["lat", "lon"],
|
| 177 |
-
},
|
| 178 |
-
),
|
| 179 |
-
]
|
| 180 |
-
|
| 181 |
-
_ACTION_BY_TOOL: dict[str, Callable[[dict[str, Any]], Any]] = {
|
| 182 |
-
"look": lambda a: LookAction(
|
| 183 |
-
heading_deg=float(a["heading_deg"]),
|
| 184 |
-
pitch_deg=float(a.get("pitch_deg", 0.0)),
|
| 185 |
-
fov_deg=float(a.get("fov_deg", 90.0)),
|
| 186 |
-
),
|
| 187 |
-
"pan": lambda a: PanAction(delta_deg=float(a["delta_deg"])),
|
| 188 |
-
"zoom": lambda a: ZoomAction(fov_deg=float(a["fov_deg"])),
|
| 189 |
-
"move": lambda a: MoveAction(
|
| 190 |
-
direction=str(a["direction"]), meters=float(a.get("meters", 10.0))
|
| 191 |
-
),
|
| 192 |
-
"place_pin": lambda a: PinAction(
|
| 193 |
-
lat=float(a["lat"]),
|
| 194 |
-
lon=float(a["lon"]),
|
| 195 |
-
label=a.get("label") or None,
|
| 196 |
-
span_deg=float(a.get("span_deg", 7.0)),
|
| 197 |
-
),
|
| 198 |
-
"view_map": lambda a: ViewMapAction(
|
| 199 |
-
lat=float(a["lat"]),
|
| 200 |
-
lon=float(a["lon"]),
|
| 201 |
-
span_deg=float(a.get("span_deg", 7.0)),
|
| 202 |
-
),
|
| 203 |
-
"measure": lambda a: MeasureAction(
|
| 204 |
-
lat_a=float(a["lat_a"]),
|
| 205 |
-
lon_a=float(a["lon_a"]),
|
| 206 |
-
lat_b=float(a["lat_b"]),
|
| 207 |
-
lon_b=float(a["lon_b"]),
|
| 208 |
-
),
|
| 209 |
-
"submit_guess": lambda a: GuessAction(
|
| 210 |
-
lat=float(a["lat"]),
|
| 211 |
-
lon=float(a["lon"]),
|
| 212 |
-
country=a.get("country") or None,
|
| 213 |
-
confidence=(
|
| 214 |
-
float(a["confidence"]) if a.get("confidence") is not None else None
|
| 215 |
-
),
|
| 216 |
-
reasoning=a.get("reasoning") or None,
|
| 217 |
-
),
|
| 218 |
-
}
|
| 219 |
-
|
| 220 |
-
|
| 221 |
-
def load_tasks(
|
| 222 |
-
index_path: str | pathlib.Path,
|
| 223 |
-
repeat: int = 1,
|
| 224 |
-
split: str | None = None,
|
| 225 |
-
) -> list[dict[str, Any]]:
|
| 226 |
-
"""
|
| 227 |
-
Read a frozen task index into harness task dicts.
|
| 228 |
-
|
| 229 |
-
Args:
|
| 230 |
-
index_path (`str` or `pathlib.Path`):
|
| 231 |
-
The JSONL index written by `scripts/build_tasks.py`.
|
| 232 |
-
repeat (`int`, *optional*, defaults to `1`):
|
| 233 |
-
Emit each task this many times consecutively. `repeat=16` gives a
|
| 234 |
-
GRPO group of 16 rollouts per location.
|
| 235 |
-
split (`str`, *optional*):
|
| 236 |
-
Split name to record on every task, so the session factory selects
|
| 237 |
-
from the right index server-side. Required whenever the server
|
| 238 |
-
serves more than one split, because an index alone does not say
|
| 239 |
-
which one it is.
|
| 240 |
-
|
| 241 |
-
Returns:
|
| 242 |
-
`list[dict]`: One dict per episode, each with `task_index`, `task_id`,
|
| 243 |
-
`country` and, when given, `split`.
|
| 244 |
-
|
| 245 |
-
Examples:
|
| 246 |
-
|
| 247 |
-
```python
|
| 248 |
-
tasks = load_tasks("tasks/eval_pano_v3.jsonl", split="eval")
|
| 249 |
-
tasks = load_tasks("tasks/train_pano_v3.jsonl", repeat=16, split="train")
|
| 250 |
-
```
|
| 251 |
-
"""
|
| 252 |
-
rows = []
|
| 253 |
-
for line in pathlib.Path(index_path).read_text().splitlines():
|
| 254 |
-
if not line.strip():
|
| 255 |
-
continue
|
| 256 |
-
row = json.loads(line)
|
| 257 |
-
task = {
|
| 258 |
-
"task_index": int(row["task_index"]),
|
| 259 |
-
"task_id": str(row["task_id"]),
|
| 260 |
-
"country": str(row.get("country", "")),
|
| 261 |
-
}
|
| 262 |
-
if split is not None:
|
| 263 |
-
task["split"] = split
|
| 264 |
-
rows.extend([dict(task) for _ in range(repeat)])
|
| 265 |
-
return rows
|
| 266 |
-
|
| 267 |
-
|
| 268 |
-
def cycle_tasks(tasks: list[dict[str, Any]]) -> Iterator[dict[str, Any]]:
|
| 269 |
-
"""Yield tasks forever, so `num_episodes` may exceed the index size."""
|
| 270 |
-
while True:
|
| 271 |
-
for task in tasks:
|
| 272 |
-
yield dict(task)
|
| 273 |
-
|
| 274 |
-
|
| 275 |
-
def _initial_messages(result: Any, task: Any) -> list[dict[str, Any]]:
|
| 276 |
-
"""Opening message: the prompt plus the first view, as an image part."""
|
| 277 |
-
observation = result.observation
|
| 278 |
-
content: list[dict[str, Any]] = [{"type": "text", "text": observation.prompt}]
|
| 279 |
-
if observation.image_base64:
|
| 280 |
-
content.append(
|
| 281 |
-
{
|
| 282 |
-
"type": "image_url",
|
| 283 |
-
"image_url": {
|
| 284 |
-
"url": f"data:image/jpeg;base64,{observation.image_base64}"
|
| 285 |
-
},
|
| 286 |
-
}
|
| 287 |
-
)
|
| 288 |
-
return [{"role": "user", "content": content}]
|
| 289 |
-
|
| 290 |
-
|
| 291 |
-
def _tool_result(
|
| 292 |
-
tool_name: str, arguments: dict[str, Any], result: Any, state: Any
|
| 293 |
-
) -> ToolResult:
|
| 294 |
-
"""Feed the observation back as text plus, when present, an image."""
|
| 295 |
-
observation = result.observation
|
| 296 |
-
data: dict[str, Any] = {
|
| 297 |
-
"feedback": observation.feedback,
|
| 298 |
-
"heading_deg": observation.heading_deg,
|
| 299 |
-
"fov_deg": observation.fov_deg,
|
| 300 |
-
"steps_remaining": observation.steps_remaining,
|
| 301 |
-
"pins": [p.model_dump() for p in observation.pins],
|
| 302 |
-
}
|
| 303 |
-
if observation.image_base64:
|
| 304 |
-
data["image_base64"] = observation.image_base64
|
| 305 |
-
data["image_kind"] = observation.image_kind
|
| 306 |
-
if observation.distance_km is not None:
|
| 307 |
-
data["distance_km"] = observation.distance_km
|
| 308 |
-
return ToolResult(
|
| 309 |
-
data=data,
|
| 310 |
-
done=bool(result.done),
|
| 311 |
-
metadata={
|
| 312 |
-
"reward": result.reward,
|
| 313 |
-
"tool": tool_name,
|
| 314 |
-
"state": state.model_dump() if hasattr(state, "model_dump") else state,
|
| 315 |
-
},
|
| 316 |
-
)
|
| 317 |
-
|
| 318 |
-
|
| 319 |
-
def _verify(
|
| 320 |
-
transcript: list[dict[str, Any]],
|
| 321 |
-
final_state: Any,
|
| 322 |
-
last_result: Any,
|
| 323 |
-
task: Any,
|
| 324 |
-
) -> VerifyResult:
|
| 325 |
-
"""
|
| 326 |
-
Summarise a finished episode for the collector.
|
| 327 |
-
|
| 328 |
-
`env_reward` forwards the reward the environment itself computed. Domain
|
| 329 |
-
knowledge belongs inside the environment, so nothing here recomputes or
|
| 330 |
-
adjusts it; the extra fields are derived statistics only.
|
| 331 |
-
"""
|
| 332 |
-
observation = getattr(last_result, "observation", None)
|
| 333 |
-
distance = getattr(observation, "distance_km", None)
|
| 334 |
-
env_reward = getattr(last_result, "reward", None)
|
| 335 |
-
metrics: dict[str, Any] = {
|
| 336 |
-
"distance_km": float(distance) if distance is not None else -1.0,
|
| 337 |
-
"parsed_ok": float(bool(getattr(observation, "parsed_ok", False))),
|
| 338 |
-
"guessed": float(distance is not None),
|
| 339 |
-
"within_200km": float(distance is not None and distance < 200.0),
|
| 340 |
-
"within_25km": float(distance is not None and distance < 25.0),
|
| 341 |
-
}
|
| 342 |
-
if isinstance(task, dict) and task.get("task_index") is not None:
|
| 343 |
-
metrics["task_index"] = float(task["task_index"])
|
| 344 |
-
return VerifyResult(
|
| 345 |
-
env_reward=float(env_reward) if env_reward is not None else None,
|
| 346 |
-
done=bool(getattr(last_result, "done", False)),
|
| 347 |
-
metrics=metrics,
|
| 348 |
-
)
|
| 349 |
-
|
| 350 |
-
|
| 351 |
-
class GeoGuesserSessionFactory(ResourceSessionFactory):
|
| 352 |
-
"""Create GeoGuesser-backed resource sessions for harness rollouts.
|
| 353 |
-
|
| 354 |
-
Args:
|
| 355 |
-
client_factory (`Callable[[], GeoGuesserEnv]`):
|
| 356 |
-
Builds a client per session, so concurrent sessions stay isolated.
|
| 357 |
-
include_navigation (`bool`, *optional*, defaults to `True`):
|
| 358 |
-
Advertise `move` in the tool list. Set `False` for a task index
|
| 359 |
-
whose backend cannot walk, so the model never sees a dead tool.
|
| 360 |
-
|
| 361 |
-
Examples:
|
| 362 |
-
|
| 363 |
-
```python
|
| 364 |
-
factory = GeoGuesserSessionFactory(
|
| 365 |
-
lambda: GeoGuesserEnv(base_url="http://localhost:8000")
|
| 366 |
-
)
|
| 367 |
-
session = factory.create(task={"task_index": 7})
|
| 368 |
-
```
|
| 369 |
-
"""
|
| 370 |
-
|
| 371 |
-
def __init__(
|
| 372 |
-
self,
|
| 373 |
-
client_factory: Callable[[], GeoGuesserEnv],
|
| 374 |
-
*,
|
| 375 |
-
include_navigation: bool = True,
|
| 376 |
-
):
|
| 377 |
-
self._client_factory = client_factory
|
| 378 |
-
self._tools = [
|
| 379 |
-
tool
|
| 380 |
-
for tool in GEOGUESSER_TOOLS
|
| 381 |
-
if include_navigation or tool.name != "move"
|
| 382 |
-
]
|
| 383 |
-
|
| 384 |
-
def create(
|
| 385 |
-
self,
|
| 386 |
-
task: Any = None,
|
| 387 |
-
seed: int | None = None,
|
| 388 |
-
episode_id: str | None = None,
|
| 389 |
-
) -> StepEnvSessionAdapter:
|
| 390 |
-
"""
|
| 391 |
-
Open a session on one task.
|
| 392 |
-
|
| 393 |
-
The task's `split` and `task_index` are passed through `reset_kwargs`,
|
| 394 |
-
which is what makes a rollout reproducible: the same task always starts
|
| 395 |
-
from the same panorama at the same heading. Without the split, an index
|
| 396 |
-
is ambiguous once the server serves more than one.
|
| 397 |
-
|
| 398 |
-
Args:
|
| 399 |
-
task (`dict`, *optional*):
|
| 400 |
-
Task dict with a `task_index` and optionally a `split`. `None`
|
| 401 |
-
selects randomly from the server's default split.
|
| 402 |
-
seed (`int`, *optional*):
|
| 403 |
-
Fallback selector when no task is given.
|
| 404 |
-
episode_id (`str`, *optional*):
|
| 405 |
-
Episode identifier recorded by the collector.
|
| 406 |
-
|
| 407 |
-
Returns:
|
| 408 |
-
`StepEnvSessionAdapter`: The session, ready to be driven.
|
| 409 |
-
"""
|
| 410 |
-
reset_kwargs: dict[str, Any] = {}
|
| 411 |
-
if isinstance(task, dict):
|
| 412 |
-
if task.get("split"):
|
| 413 |
-
reset_kwargs["split"] = str(task["split"])
|
| 414 |
-
if task.get("task_index") is not None:
|
| 415 |
-
reset_kwargs["index"] = int(task["task_index"])
|
| 416 |
-
elif isinstance(task, int):
|
| 417 |
-
reset_kwargs["index"] = task
|
| 418 |
-
|
| 419 |
-
return StepEnvSessionAdapter(
|
| 420 |
-
client=self._client_factory(),
|
| 421 |
-
task=task,
|
| 422 |
-
seed=seed,
|
| 423 |
-
episode_id=episode_id,
|
| 424 |
-
tool_specs=list(self._tools),
|
| 425 |
-
action_builder=lambda name, arguments: to_wire(
|
| 426 |
-
_ACTION_BY_TOOL[name](arguments)
|
| 427 |
-
),
|
| 428 |
-
initial_messages_builder=_initial_messages,
|
| 429 |
-
tool_result_builder=_tool_result,
|
| 430 |
-
verify_builder=_verify,
|
| 431 |
-
reset_kwargs=reset_kwargs,
|
| 432 |
-
)
|
| 433 |
-
|
| 434 |
-
|
| 435 |
-
__all__ = [
|
| 436 |
-
"GEOGUESSER_TOOLS",
|
| 437 |
-
"GeoGuesserSessionFactory",
|
| 438 |
-
"cycle_tasks",
|
| 439 |
-
"load_tasks",
|
| 440 |
-
]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/models.py
DELETED
|
@@ -1,462 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Data models for the GeoGuesser environment.
|
| 4 |
-
|
| 5 |
-
An episode places the agent at an unknown street-level location. It may look
|
| 6 |
-
around, walk along the road, and pin candidate coordinates on a map before
|
| 7 |
-
committing to a final guess. Only the final guess is scored.
|
| 8 |
-
"""
|
| 9 |
-
|
| 10 |
-
from __future__ import annotations
|
| 11 |
-
|
| 12 |
-
from enum import Enum
|
| 13 |
-
from typing import Any, Literal
|
| 14 |
-
|
| 15 |
-
from openenv.core.env_server import Action, Observation, State
|
| 16 |
-
from pydantic import BaseModel, Field
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
class EpisodeMode(str, Enum):
|
| 20 |
-
"""How much of the tool surface an episode exposes.
|
| 21 |
-
|
| 22 |
-
Attributes:
|
| 23 |
-
SINGLE_SHOT:
|
| 24 |
-
One view, one guess. No pin loop, no navigation.
|
| 25 |
-
AGENTIC:
|
| 26 |
-
The full tool surface, subject to what the backend supports.
|
| 27 |
-
NMPZ:
|
| 28 |
-
No move, pan or zoom — the competitive "NMPZ" mode. Only pinning
|
| 29 |
-
and guessing.
|
| 30 |
-
"""
|
| 31 |
-
|
| 32 |
-
SINGLE_SHOT = "single_shot"
|
| 33 |
-
AGENTIC = "agentic"
|
| 34 |
-
NMPZ = "nmpz"
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
class RewardMode(str, Enum):
|
| 38 |
-
"""Which ground-truth granularity the reward is computed against."""
|
| 39 |
-
|
| 40 |
-
COORDS = "coords"
|
| 41 |
-
COUNTRY_ONLY = "country_only"
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
# =============================================================================
|
| 45 |
-
# Actions
|
| 46 |
-
# =============================================================================
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
class TypedAction(Action):
|
| 50 |
-
"""Base class for the ergonomic, per-operation action types.
|
| 51 |
-
|
| 52 |
-
These are what callers construct in Python. They are converted to the flat
|
| 53 |
-
[`GeoGuesserAction`] wire type by the client, because an HTTP environment
|
| 54 |
-
declares exactly one action schema and `Action` forbids unknown fields.
|
| 55 |
-
"""
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
class LookAction(TypedAction):
|
| 59 |
-
"""Render a perspective view out of the current panorama.
|
| 60 |
-
|
| 61 |
-
Args:
|
| 62 |
-
heading_deg (`float`):
|
| 63 |
-
Absolute compass heading in degrees, `0` being true north.
|
| 64 |
-
pitch_deg (`float`, *optional*, defaults to `0.0`):
|
| 65 |
-
Vertical angle in degrees; positive looks up.
|
| 66 |
-
fov_deg (`float`, *optional*, defaults to `90.0`):
|
| 67 |
-
Horizontal field of view. Smaller values zoom in.
|
| 68 |
-
"""
|
| 69 |
-
|
| 70 |
-
heading_deg: float = Field(default=0.0, ge=-3600.0, le=3600.0)
|
| 71 |
-
pitch_deg: float = Field(default=0.0, ge=-90.0, le=90.0)
|
| 72 |
-
fov_deg: float = Field(default=90.0, ge=10.0, le=120.0)
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
class PanAction(TypedAction):
|
| 76 |
-
"""Turn relative to the current heading.
|
| 77 |
-
|
| 78 |
-
Args:
|
| 79 |
-
delta_deg (`float`):
|
| 80 |
-
Degrees to turn; positive turns right.
|
| 81 |
-
"""
|
| 82 |
-
|
| 83 |
-
delta_deg: float = Field(ge=-3600.0, le=3600.0)
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
class ZoomAction(TypedAction):
|
| 87 |
-
"""Change the field of view without turning.
|
| 88 |
-
|
| 89 |
-
Args:
|
| 90 |
-
fov_deg (`float`):
|
| 91 |
-
New horizontal field of view in degrees.
|
| 92 |
-
"""
|
| 93 |
-
|
| 94 |
-
fov_deg: float = Field(ge=10.0, le=120.0)
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
class MoveAction(TypedAction):
|
| 98 |
-
"""Walk along the captured sequence.
|
| 99 |
-
|
| 100 |
-
Args:
|
| 101 |
-
direction (`str`):
|
| 102 |
-
Either `"forward"` or `"backward"` along the sequence.
|
| 103 |
-
meters (`float`, *optional*, defaults to `10.0`):
|
| 104 |
-
Requested distance. Frame spacing is irregular, so the observation
|
| 105 |
-
reports how far the move actually travelled.
|
| 106 |
-
"""
|
| 107 |
-
|
| 108 |
-
direction: str = Field(pattern="^(forward|backward)$")
|
| 109 |
-
meters: float = Field(default=10.0, gt=0.0, le=500.0)
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
class PinAction(TypedAction):
|
| 113 |
-
"""Place a candidate pin and receive a map of where it landed.
|
| 114 |
-
|
| 115 |
-
The response describes the pinned location only. It carries no information
|
| 116 |
-
about the true location.
|
| 117 |
-
|
| 118 |
-
Args:
|
| 119 |
-
lat (`float`):
|
| 120 |
-
Latitude of the candidate.
|
| 121 |
-
lon (`float`):
|
| 122 |
-
Longitude of the candidate.
|
| 123 |
-
label (`str`, *optional*):
|
| 124 |
-
Free-text note carried back in the pin list.
|
| 125 |
-
span_deg (`float`, *optional*, defaults to `7.0`):
|
| 126 |
-
Half-width in degrees of the map window returned with the pin.
|
| 127 |
-
Choosing the zoom matters: below roughly 4 degrees the map adds
|
| 128 |
-
roads, urban areas and town names, which is what makes aiming
|
| 129 |
-
within a city possible rather than guessing at its centre.
|
| 130 |
-
"""
|
| 131 |
-
|
| 132 |
-
lat: float = Field(ge=-90.0, le=90.0)
|
| 133 |
-
lon: float = Field(ge=-180.0, le=180.0)
|
| 134 |
-
label: str | None = None
|
| 135 |
-
span_deg: float = Field(default=7.0, gt=0.02, le=180.0)
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
class ViewMapAction(TypedAction):
|
| 139 |
-
"""Pan and zoom the map without committing a pin.
|
| 140 |
-
|
| 141 |
-
Args:
|
| 142 |
-
lat (`float`):
|
| 143 |
-
Latitude at the centre of the view.
|
| 144 |
-
lon (`float`):
|
| 145 |
-
Longitude at the centre of the view.
|
| 146 |
-
span_deg (`float`, *optional*, defaults to `7.0`):
|
| 147 |
-
Half-width of the window in degrees.
|
| 148 |
-
"""
|
| 149 |
-
|
| 150 |
-
lat: float = Field(ge=-90.0, le=90.0)
|
| 151 |
-
lon: float = Field(ge=-180.0, le=180.0)
|
| 152 |
-
span_deg: float = Field(default=7.0, gt=0.05, le=180.0)
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
class MeasureAction(TypedAction):
|
| 156 |
-
"""Great-circle distance between two of the agent's own coordinates."""
|
| 157 |
-
|
| 158 |
-
lat_a: float = Field(ge=-90.0, le=90.0)
|
| 159 |
-
lon_a: float = Field(ge=-180.0, le=180.0)
|
| 160 |
-
lat_b: float = Field(ge=-90.0, le=90.0)
|
| 161 |
-
lon_b: float = Field(ge=-180.0, le=180.0)
|
| 162 |
-
|
| 163 |
-
|
| 164 |
-
class GuessAction(TypedAction):
|
| 165 |
-
"""Commit a final guess. Terminal.
|
| 166 |
-
|
| 167 |
-
Either supply `response` and let the environment parse it, or supply
|
| 168 |
-
`lat`/`lon` directly. Passing the raw reply keeps extraction failures
|
| 169 |
-
visible in the score rather than hidden in the harness.
|
| 170 |
-
|
| 171 |
-
Args:
|
| 172 |
-
response (`str`, *optional*):
|
| 173 |
-
The model's unedited reply. Coordinates are extracted from it.
|
| 174 |
-
lat (`float`, *optional*):
|
| 175 |
-
Latitude, when the caller has already parsed the reply.
|
| 176 |
-
lon (`float`, *optional*):
|
| 177 |
-
Longitude, when the caller has already parsed the reply.
|
| 178 |
-
country (`str`, *optional*):
|
| 179 |
-
ISO-3166 alpha-2 code or country name, scored for partial credit.
|
| 180 |
-
confidence (`float`, *optional*):
|
| 181 |
-
Self-reported confidence in [0, 1], recorded for calibration.
|
| 182 |
-
reasoning (`str`, *optional*):
|
| 183 |
-
Free-text rationale, recorded but not scored.
|
| 184 |
-
"""
|
| 185 |
-
|
| 186 |
-
response: str | None = None
|
| 187 |
-
lat: float | None = Field(default=None, ge=-90.0, le=90.0)
|
| 188 |
-
lon: float | None = Field(default=None, ge=-180.0, le=180.0)
|
| 189 |
-
country: str | None = None
|
| 190 |
-
confidence: float | None = Field(default=None, ge=0.0, le=1.0)
|
| 191 |
-
reasoning: str | None = None
|
| 192 |
-
|
| 193 |
-
|
| 194 |
-
# =============================================================================
|
| 195 |
-
# Wire action
|
| 196 |
-
# =============================================================================
|
| 197 |
-
|
| 198 |
-
|
| 199 |
-
class GeoGuesserAction(Action):
|
| 200 |
-
"""The single action schema the server accepts.
|
| 201 |
-
|
| 202 |
-
An HTTP environment declares one action class, so every operation travels
|
| 203 |
-
as this flat record with an `op` discriminator. Callers normally build a
|
| 204 |
-
[`TypedAction`] subclass instead and let the client convert.
|
| 205 |
-
|
| 206 |
-
Attributes:
|
| 207 |
-
op (`str`):
|
| 208 |
-
Which operation to perform: `"look"`, `"pan"`, `"zoom"`, `"move"`,
|
| 209 |
-
`"pin"`, `"view_map"`, `"measure"` or `"guess"`.
|
| 210 |
-
"""
|
| 211 |
-
|
| 212 |
-
op: Literal["look", "pan", "zoom", "move", "pin", "view_map", "measure", "guess"]
|
| 213 |
-
|
| 214 |
-
heading_deg: float | None = None
|
| 215 |
-
pitch_deg: float | None = None
|
| 216 |
-
fov_deg: float | None = None
|
| 217 |
-
delta_deg: float | None = None
|
| 218 |
-
direction: str | None = None
|
| 219 |
-
meters: float | None = None
|
| 220 |
-
lat: float | None = None
|
| 221 |
-
lon: float | None = None
|
| 222 |
-
label: str | None = None
|
| 223 |
-
span_deg: float | None = None
|
| 224 |
-
lat_a: float | None = None
|
| 225 |
-
lon_a: float | None = None
|
| 226 |
-
lat_b: float | None = None
|
| 227 |
-
lon_b: float | None = None
|
| 228 |
-
response: str | None = None
|
| 229 |
-
country: str | None = None
|
| 230 |
-
confidence: float | None = None
|
| 231 |
-
reasoning: str | None = None
|
| 232 |
-
|
| 233 |
-
|
| 234 |
-
_OP_BY_TYPE: dict[type, str] = {}
|
| 235 |
-
_TYPE_BY_OP: dict[str, type] = {}
|
| 236 |
-
|
| 237 |
-
|
| 238 |
-
def _register_ops() -> None:
|
| 239 |
-
pairs = [
|
| 240 |
-
(LookAction, "look"),
|
| 241 |
-
(PanAction, "pan"),
|
| 242 |
-
(ZoomAction, "zoom"),
|
| 243 |
-
(MoveAction, "move"),
|
| 244 |
-
(PinAction, "pin"),
|
| 245 |
-
(ViewMapAction, "view_map"),
|
| 246 |
-
(MeasureAction, "measure"),
|
| 247 |
-
(GuessAction, "guess"),
|
| 248 |
-
]
|
| 249 |
-
for cls, op in pairs:
|
| 250 |
-
_OP_BY_TYPE[cls] = op
|
| 251 |
-
_TYPE_BY_OP[op] = cls
|
| 252 |
-
|
| 253 |
-
|
| 254 |
-
_register_ops()
|
| 255 |
-
|
| 256 |
-
|
| 257 |
-
def to_wire(action: TypedAction) -> GeoGuesserAction:
|
| 258 |
-
"""
|
| 259 |
-
Convert a typed action into the flat wire action.
|
| 260 |
-
|
| 261 |
-
Args:
|
| 262 |
-
action ([`TypedAction`]):
|
| 263 |
-
The action to convert.
|
| 264 |
-
|
| 265 |
-
Returns:
|
| 266 |
-
[`GeoGuesserAction`]: The same action, flattened, with `op` set.
|
| 267 |
-
"""
|
| 268 |
-
op = _OP_BY_TYPE.get(type(action))
|
| 269 |
-
if op is None:
|
| 270 |
-
raise TypeError(f"No wire op registered for {type(action).__name__}")
|
| 271 |
-
payload = action.model_dump(exclude_none=True, exclude={"metadata"})
|
| 272 |
-
return GeoGuesserAction(op=op, **payload)
|
| 273 |
-
|
| 274 |
-
|
| 275 |
-
def from_wire(action: GeoGuesserAction) -> TypedAction:
|
| 276 |
-
"""
|
| 277 |
-
Rebuild the typed action a wire action stands for.
|
| 278 |
-
|
| 279 |
-
Args:
|
| 280 |
-
action ([`GeoGuesserAction`]):
|
| 281 |
-
The received wire action.
|
| 282 |
-
|
| 283 |
-
Returns:
|
| 284 |
-
[`TypedAction`]: The corresponding typed action, validated.
|
| 285 |
-
"""
|
| 286 |
-
cls = _TYPE_BY_OP.get(action.op)
|
| 287 |
-
if cls is None:
|
| 288 |
-
raise ValueError(f"Unknown op: {action.op!r}")
|
| 289 |
-
fields = set(cls.model_fields) - {"metadata"}
|
| 290 |
-
payload = {
|
| 291 |
-
k: v for k, v in action.model_dump(exclude_none=True).items() if k in fields
|
| 292 |
-
}
|
| 293 |
-
return cls(**payload)
|
| 294 |
-
|
| 295 |
-
|
| 296 |
-
# =============================================================================
|
| 297 |
-
# Observation
|
| 298 |
-
# =============================================================================
|
| 299 |
-
|
| 300 |
-
|
| 301 |
-
class Pin(BaseModel):
|
| 302 |
-
"""One candidate pin and what the environment could say about it.
|
| 303 |
-
|
| 304 |
-
Attributes:
|
| 305 |
-
index (`int`):
|
| 306 |
-
1-based position in the pin list.
|
| 307 |
-
lat (`float`):
|
| 308 |
-
Latitude of the candidate.
|
| 309 |
-
lon (`float`):
|
| 310 |
-
Longitude of the candidate.
|
| 311 |
-
label (`str` or `None`):
|
| 312 |
-
The note the agent attached, if any.
|
| 313 |
-
description (`str`):
|
| 314 |
-
What the environment could say about the pinned coordinate. Never
|
| 315 |
-
anything about the target.
|
| 316 |
-
"""
|
| 317 |
-
|
| 318 |
-
index: int
|
| 319 |
-
lat: float
|
| 320 |
-
lon: float
|
| 321 |
-
label: str | None = None
|
| 322 |
-
description: str = ""
|
| 323 |
-
|
| 324 |
-
|
| 325 |
-
class GeoGuesserObservation(Observation):
|
| 326 |
-
"""What the agent sees after a reset or a step.
|
| 327 |
-
|
| 328 |
-
The schema is identical across backends. A capability the backend lacks
|
| 329 |
-
shows up as an unregistered tool and an empty field, never as a different
|
| 330 |
-
shape, so one policy runs against every backend.
|
| 331 |
-
|
| 332 |
-
Attributes:
|
| 333 |
-
prompt (`str`):
|
| 334 |
-
Instructions, populated on reset.
|
| 335 |
-
image_base64 (`str` or `None`):
|
| 336 |
-
The most recent rendered image as base64 PNG or JPEG — a
|
| 337 |
-
perspective view, or a map after a pin.
|
| 338 |
-
image_kind (`str`):
|
| 339 |
-
Either `"view"`, `"map"` or `"none"`, saying what the image shows.
|
| 340 |
-
heading_deg (`float`):
|
| 341 |
-
Current compass heading in degrees.
|
| 342 |
-
pitch_deg (`float`):
|
| 343 |
-
Current vertical angle in degrees.
|
| 344 |
-
fov_deg (`float`):
|
| 345 |
-
Current field of view in degrees.
|
| 346 |
-
moved_meters (`float`):
|
| 347 |
-
Distance actually travelled by the last move.
|
| 348 |
-
total_moved_meters (`float`):
|
| 349 |
-
Cumulative distance travelled this episode.
|
| 350 |
-
can_move_forward (`bool`):
|
| 351 |
-
Whether a forward frame exists on the sequence.
|
| 352 |
-
can_move_backward (`bool`):
|
| 353 |
-
Whether a backward frame exists on the sequence.
|
| 354 |
-
available_tools (`list[str]`):
|
| 355 |
-
Tool names this backend actually registered.
|
| 356 |
-
steps_remaining (`int`):
|
| 357 |
-
Actions left before the episode is cut off.
|
| 358 |
-
pins (`list[Pin]`):
|
| 359 |
-
Candidates placed so far, in order.
|
| 360 |
-
feedback (`str`):
|
| 361 |
-
Text describing the result of the last action. For a pin, this
|
| 362 |
-
describes the pinned location and nothing about the target.
|
| 363 |
-
captured_at (`str`):
|
| 364 |
-
Capture date of the current panorama, `YYYY-MM` — a legitimate
|
| 365 |
-
meta clue, as in the real game.
|
| 366 |
-
distance_km (`float` or `None`):
|
| 367 |
-
Distance from guess to truth. Populated only after a guess.
|
| 368 |
-
score (`float` or `None`):
|
| 369 |
-
Distance score in [0, 1], before action costs. After a guess only.
|
| 370 |
-
action_cost (`float`):
|
| 371 |
-
Reward already spent on information gathering this episode. Visible
|
| 372 |
-
throughout, not only after the guess, so a policy can see what it
|
| 373 |
-
has committed.
|
| 374 |
-
true_lat (`float` or `None`):
|
| 375 |
-
Ground truth latitude, revealed only after a guess.
|
| 376 |
-
true_lon (`float` or `None`):
|
| 377 |
-
Ground truth longitude, revealed only after a guess.
|
| 378 |
-
parsed_ok (`bool`):
|
| 379 |
-
Whether coordinates could be extracted from the guess.
|
| 380 |
-
"""
|
| 381 |
-
|
| 382 |
-
prompt: str = ""
|
| 383 |
-
image_base64: str | None = None
|
| 384 |
-
image_kind: str = "none"
|
| 385 |
-
|
| 386 |
-
heading_deg: float = 0.0
|
| 387 |
-
pitch_deg: float = 0.0
|
| 388 |
-
fov_deg: float = 90.0
|
| 389 |
-
|
| 390 |
-
moved_meters: float = 0.0
|
| 391 |
-
total_moved_meters: float = 0.0
|
| 392 |
-
can_move_forward: bool = False
|
| 393 |
-
can_move_backward: bool = False
|
| 394 |
-
|
| 395 |
-
available_tools: list[str] = Field(default_factory=list)
|
| 396 |
-
steps_remaining: int = 0
|
| 397 |
-
pins: list[Pin] = Field(default_factory=list)
|
| 398 |
-
feedback: str = ""
|
| 399 |
-
captured_at: str = ""
|
| 400 |
-
|
| 401 |
-
distance_km: float | None = None
|
| 402 |
-
score: float | None = None
|
| 403 |
-
action_cost: float | None = None
|
| 404 |
-
true_lat: float | None = None
|
| 405 |
-
true_lon: float | None = None
|
| 406 |
-
parsed_ok: bool = True
|
| 407 |
-
|
| 408 |
-
|
| 409 |
-
# =============================================================================
|
| 410 |
-
# State
|
| 411 |
-
# =============================================================================
|
| 412 |
-
|
| 413 |
-
|
| 414 |
-
class GeoGuesserState(State):
|
| 415 |
-
"""Internal episode state. Never sent to the agent verbatim.
|
| 416 |
-
|
| 417 |
-
Attributes:
|
| 418 |
-
task_index (`int`):
|
| 419 |
-
Index into the frozen task list, `-1` before the first reset.
|
| 420 |
-
task_id (`str`):
|
| 421 |
-
Stable identifier of the sampled task.
|
| 422 |
-
frame_index (`int`):
|
| 423 |
-
Position within the task's sequence.
|
| 424 |
-
heading_deg (`float`):
|
| 425 |
-
Current heading in degrees.
|
| 426 |
-
pitch_deg (`float`):
|
| 427 |
-
Current pitch in degrees.
|
| 428 |
-
fov_deg (`float`):
|
| 429 |
-
Current field of view in degrees.
|
| 430 |
-
n_looks (`int`):
|
| 431 |
-
Count of view renders, for action cost.
|
| 432 |
-
n_maps (`int`):
|
| 433 |
-
Count of map renders that were not pins.
|
| 434 |
-
n_pins (`int`):
|
| 435 |
-
Count of pins placed.
|
| 436 |
-
n_moves (`int`):
|
| 437 |
-
Count of moves taken.
|
| 438 |
-
n_free (`int`):
|
| 439 |
-
Count of free-tool calls, which cost nothing but are capped so they
|
| 440 |
-
cannot be issued forever.
|
| 441 |
-
total_moved_meters (`float`):
|
| 442 |
-
Cumulative distance travelled.
|
| 443 |
-
submitted (`bool`):
|
| 444 |
-
Whether the single allowed guess has been made.
|
| 445 |
-
pins (`list[dict]`):
|
| 446 |
-
Pins placed so far.
|
| 447 |
-
"""
|
| 448 |
-
|
| 449 |
-
task_index: int = -1
|
| 450 |
-
task_id: str = ""
|
| 451 |
-
frame_index: int = 0
|
| 452 |
-
heading_deg: float = 0.0
|
| 453 |
-
pitch_deg: float = 0.0
|
| 454 |
-
fov_deg: float = 90.0
|
| 455 |
-
n_looks: int = 0
|
| 456 |
-
n_maps: int = 0
|
| 457 |
-
n_pins: int = 0
|
| 458 |
-
n_moves: int = 0
|
| 459 |
-
n_free: int = 0
|
| 460 |
-
total_moved_meters: float = 0.0
|
| 461 |
-
submitted: bool = False
|
| 462 |
-
pins: list[dict[str, Any]] = Field(default_factory=list)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/__init__.py
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Server-side implementation of the GeoGuesser environment."""
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/app.py
DELETED
|
@@ -1,349 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""FastAPI application for the GeoGuesser environment.
|
| 4 |
-
|
| 5 |
-
Configuration comes from the process environment so one image can serve every
|
| 6 |
-
variant. `MAPILLARY_API_KEY` is needed only to fill cache misses; with a warm
|
| 7 |
-
cache the server runs with no network access at all.
|
| 8 |
-
|
| 9 |
-
Usage:
|
| 10 |
-
uv run --project . server
|
| 11 |
-
uvicorn server.app:app --host 0.0.0.0 --port 8000
|
| 12 |
-
"""
|
| 13 |
-
|
| 14 |
-
from __future__ import annotations
|
| 15 |
-
|
| 16 |
-
import inspect
|
| 17 |
-
import logging
|
| 18 |
-
import os
|
| 19 |
-
import pathlib
|
| 20 |
-
|
| 21 |
-
from openenv.core.env_server.http_server import create_app
|
| 22 |
-
|
| 23 |
-
try:
|
| 24 |
-
from geoguesser_env.models import GeoGuesserAction, GeoGuesserObservation
|
| 25 |
-
from geoguesser_env.server.geoguesser_environment import GeoGuesserEnvironment
|
| 26 |
-
except ImportError: # running uvicorn from inside envs/geoguesser_env
|
| 27 |
-
from models import GeoGuesserAction, GeoGuesserObservation
|
| 28 |
-
|
| 29 |
-
from .geoguesser_environment import GeoGuesserEnvironment
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
logger = logging.getLogger(__name__)
|
| 33 |
-
|
| 34 |
-
_ROOT = pathlib.Path(__file__).resolve().parents[1]
|
| 35 |
-
|
| 36 |
-
INDEX_PATH = os.getenv("GEOGUESSER_INDEX", str(_ROOT / "tasks" / "pano_v1.jsonl"))
|
| 37 |
-
CACHE_DIR = os.getenv("GEOGUESSER_CACHE", str(_ROOT / "data" / "panos"))
|
| 38 |
-
# Named splits, as (environment variable, repo-relative default). Each is
|
| 39 |
-
# optional and a split whose file is absent is simply not offered, so the same
|
| 40 |
-
# image serves a checkout with the indexes committed, a Space with a bucket
|
| 41 |
-
# mounted at /data, and a deployment carrying only the eval set.
|
| 42 |
-
SPLIT_SOURCES = {
|
| 43 |
-
"train": ("GEOGUESSER_TASKS_TRAIN", "tasks/train_pano_v3.jsonl"),
|
| 44 |
-
"eval": ("GEOGUESSER_TASKS_EVAL", "tasks/eval_pano_v3.jsonl"),
|
| 45 |
-
}
|
| 46 |
-
DEFAULT_SPLIT = os.getenv("GEOGUESSER_DEFAULT_SPLIT", "")
|
| 47 |
-
EPISODE_MODE = os.getenv("GEOGUESSER_EPISODE_MODE", "agentic")
|
| 48 |
-
REWARD_MODE = os.getenv("GEOGUESSER_REWARD_MODE", "coords")
|
| 49 |
-
MAX_STEPS = int(os.getenv("GEOGUESSER_MAX_STEPS", "24"))
|
| 50 |
-
VIEW_SIZE = int(os.getenv("GEOGUESSER_VIEW_SIZE", "640"))
|
| 51 |
-
HIERARCHICAL = os.getenv("GEOGUESSER_HIERARCHICAL", "0") in {"1", "true", "True"}
|
| 52 |
-
# Training defaults differ from play defaults on purpose; see the reward-shape
|
| 53 |
-
# note in the README. A hosted Space serves play traffic, so it keeps the game
|
| 54 |
-
# curve and shows the licence credit unless these are set explicitly.
|
| 55 |
-
REWARD_SHAPE = os.getenv("GEOGUESSER_REWARD_SHAPE", "geoguessr")
|
| 56 |
-
COST_MODE = os.getenv("GEOGUESSER_COST_MODE", "subtract")
|
| 57 |
-
HIDE_IDENTITY = os.getenv("GEOGUESSER_HIDE_IDENTITY", "0") in {"1", "true", "True"}
|
| 58 |
-
ALLOW_FETCH = os.getenv("GEOGUESSER_ALLOW_FETCH", "1") in {"1", "true", "True"}
|
| 59 |
-
HIRES_ZOOM = os.getenv("GEOGUESSER_HIRES_ZOOM", "1") in {"1", "true", "True"}
|
| 60 |
-
REVEAL_MAP = os.getenv("GEOGUESSER_REVEAL_MAP", "1") in {"1", "true", "True"}
|
| 61 |
-
STREET_DETAIL = os.getenv("GEOGUESSER_STREET_DETAIL", "1") in {"1", "true", "True"}
|
| 62 |
-
MAX_CONCURRENT = int(os.getenv("MAX_CONCURRENT_ENVS", "4"))
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
try:
|
| 66 |
-
from .render.minimap import set_street_detail
|
| 67 |
-
except ImportError: # running uvicorn from inside envs/geoguesser_env
|
| 68 |
-
from render.minimap import set_street_detail
|
| 69 |
-
|
| 70 |
-
# Deliberately independent of ALLOW_FETCH. That flag governs Mapillary imagery,
|
| 71 |
-
# which a mirrored dataset must never reach for; street detail comes from
|
| 72 |
-
# Overpass and caches to the container's own writable disk, so a fully mirrored
|
| 73 |
-
# deployment can still draw labelled streets. Coupling the two silently gave a
|
| 74 |
-
# Space unlabelled agent maps while a local run had labelled ones.
|
| 75 |
-
set_street_detail(STREET_DETAIL)
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
def resolve_splits() -> tuple[dict[str, str], str]:
|
| 79 |
-
"""
|
| 80 |
-
Work out which named splits this deployment actually serves.
|
| 81 |
-
|
| 82 |
-
A split is offered only when its environment variable is set *and* the file
|
| 83 |
-
exists, because a Space that mounts a bucket read-only should degrade to the
|
| 84 |
-
splits it really has rather than failing to start. When none are configured
|
| 85 |
-
the legacy single `GEOGUESSER_INDEX` becomes one `train` split, so existing
|
| 86 |
-
containers behave exactly as before.
|
| 87 |
-
|
| 88 |
-
Returns:
|
| 89 |
-
`tuple` of:
|
| 90 |
-
- `dict[str, str]`: Split name to index path.
|
| 91 |
-
- `str`: Name of the default split.
|
| 92 |
-
"""
|
| 93 |
-
splits: dict[str, str] = {}
|
| 94 |
-
for name, (variable, relative) in SPLIT_SOURCES.items():
|
| 95 |
-
override = os.getenv(variable)
|
| 96 |
-
path = override or str(_ROOT / relative)
|
| 97 |
-
if not pathlib.Path(path).exists():
|
| 98 |
-
# Only complain when someone asked for it explicitly. A missing
|
| 99 |
-
# default just means this checkout has not built that split yet.
|
| 100 |
-
if override:
|
| 101 |
-
logger.warning(
|
| 102 |
-
"%s points at %s, which does not exist; "
|
| 103 |
-
"the %r split will not be offered",
|
| 104 |
-
variable,
|
| 105 |
-
path,
|
| 106 |
-
name,
|
| 107 |
-
)
|
| 108 |
-
continue
|
| 109 |
-
splits[name] = path
|
| 110 |
-
|
| 111 |
-
if not splits:
|
| 112 |
-
return {"train": INDEX_PATH}, "train"
|
| 113 |
-
|
| 114 |
-
default = DEFAULT_SPLIT or ("train" if "train" in splits else next(iter(splits)))
|
| 115 |
-
if default not in splits:
|
| 116 |
-
logger.warning(
|
| 117 |
-
"GEOGUESSER_DEFAULT_SPLIT=%r is not among %s; using %r",
|
| 118 |
-
DEFAULT_SPLIT,
|
| 119 |
-
sorted(splits),
|
| 120 |
-
next(iter(splits)),
|
| 121 |
-
)
|
| 122 |
-
default = next(iter(splits))
|
| 123 |
-
return splits, default
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
SPLITS, ACTIVE_DEFAULT_SPLIT = resolve_splits()
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
def create_geoguesser_environment() -> GeoGuesserEnvironment:
|
| 130 |
-
"""Factory: a fresh environment per WebSocket session."""
|
| 131 |
-
return GeoGuesserEnvironment(
|
| 132 |
-
splits=SPLITS,
|
| 133 |
-
default_split=ACTIVE_DEFAULT_SPLIT,
|
| 134 |
-
cache_dir=CACHE_DIR,
|
| 135 |
-
episode_mode=EPISODE_MODE,
|
| 136 |
-
max_steps=MAX_STEPS,
|
| 137 |
-
reward_mode=REWARD_MODE,
|
| 138 |
-
hierarchical_reward=HIERARCHICAL,
|
| 139 |
-
reward_shape=REWARD_SHAPE,
|
| 140 |
-
cost_mode=COST_MODE,
|
| 141 |
-
hide_task_identity=HIDE_IDENTITY,
|
| 142 |
-
view_size=VIEW_SIZE,
|
| 143 |
-
allow_fetch=ALLOW_FETCH,
|
| 144 |
-
hires_zoom=HIRES_ZOOM,
|
| 145 |
-
reveal_map=REVEAL_MAP,
|
| 146 |
-
)
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
def _build_app():
|
| 150 |
-
"""Create the app, attaching the Gradio tab when this openenv supports it."""
|
| 151 |
-
kwargs = dict(
|
| 152 |
-
env_name="geoguesser_env",
|
| 153 |
-
max_concurrent_envs=MAX_CONCURRENT,
|
| 154 |
-
)
|
| 155 |
-
signature = inspect.signature(create_app)
|
| 156 |
-
# Land people on the game, not the raw action form.
|
| 157 |
-
if "custom_tab_primary" in signature.parameters:
|
| 158 |
-
kwargs["custom_tab_primary"] = True
|
| 159 |
-
if "custom_tab_name" in signature.parameters:
|
| 160 |
-
kwargs["custom_tab_name"] = "Try Environment"
|
| 161 |
-
if "default_tab_name" in signature.parameters:
|
| 162 |
-
kwargs["default_tab_name"] = "MCP Playground"
|
| 163 |
-
if "title_override" in signature.parameters:
|
| 164 |
-
kwargs["title_override"] = "Geoguesser Environment"
|
| 165 |
-
if "gradio_builder" in signature.parameters:
|
| 166 |
-
try:
|
| 167 |
-
from .gradio_ui import build_geoguesser_gradio_app
|
| 168 |
-
|
| 169 |
-
kwargs["gradio_builder"] = build_geoguesser_gradio_app
|
| 170 |
-
except Exception as exc: # pragma: no cover - optional UI dependency
|
| 171 |
-
logger.warning("Gradio UI unavailable: %r", exc)
|
| 172 |
-
else:
|
| 173 |
-
logger.warning(
|
| 174 |
-
"Installed openenv does not support gradio_builder; "
|
| 175 |
-
"the GeoGuessr-style play tab will not be available."
|
| 176 |
-
)
|
| 177 |
-
return create_app(
|
| 178 |
-
create_geoguesser_environment,
|
| 179 |
-
GeoGuesserAction,
|
| 180 |
-
GeoGuesserObservation,
|
| 181 |
-
**kwargs,
|
| 182 |
-
)
|
| 183 |
-
|
| 184 |
-
|
| 185 |
-
def _attach_play_routes(application) -> None:
|
| 186 |
-
"""Serve panoramas and task metadata to the browser-side viewers.
|
| 187 |
-
|
| 188 |
-
The Pannellum viewer needs the raw equirectangular JPEG, which the agent
|
| 189 |
-
never receives — it only ever sees reprojected views. Ground truth is
|
| 190 |
-
exposed here because these routes exist for a human playing a round in
|
| 191 |
-
their own browser; the agent's observations still withhold it until it
|
| 192 |
-
guesses.
|
| 193 |
-
"""
|
| 194 |
-
from fastapi import HTTPException, Query
|
| 195 |
-
from fastapi.responses import FileResponse, HTMLResponse, JSONResponse
|
| 196 |
-
|
| 197 |
-
# One environment, reused for metadata only. Its per-split backends share
|
| 198 |
-
# the process-wide parsed index cache, so this is cheap.
|
| 199 |
-
catalog = create_geoguesser_environment()
|
| 200 |
-
|
| 201 |
-
def _backend(split: str | None):
|
| 202 |
-
"""Resolve a split name to its backend, as a 404 rather than a 500."""
|
| 203 |
-
try:
|
| 204 |
-
return catalog._backend_for(split or ACTIVE_DEFAULT_SPLIT)
|
| 205 |
-
except KeyError as exc:
|
| 206 |
-
raise HTTPException(status_code=404, detail=str(exc)) from exc
|
| 207 |
-
|
| 208 |
-
@application.get(
|
| 209 |
-
"/geoguesser/task/{task_index}",
|
| 210 |
-
tags=["geoguesser"],
|
| 211 |
-
response_class=JSONResponse,
|
| 212 |
-
)
|
| 213 |
-
async def geoguesser_task(task_index: int, split: str | None = Query(None)):
|
| 214 |
-
"""Metadata for one task, for the play UI."""
|
| 215 |
-
try:
|
| 216 |
-
task = _backend(split).task(task_index)
|
| 217 |
-
except IndexError as exc:
|
| 218 |
-
raise HTTPException(status_code=404, detail=str(exc)) from exc
|
| 219 |
-
frame = task.frames[task.start_frame]
|
| 220 |
-
return JSONResponse(
|
| 221 |
-
{
|
| 222 |
-
"task_index": task.task_index,
|
| 223 |
-
"task_id": task.task_id,
|
| 224 |
-
"lat": frame.lat,
|
| 225 |
-
"lon": frame.lon,
|
| 226 |
-
"country": task.country,
|
| 227 |
-
"compass_angle": frame.compass_angle,
|
| 228 |
-
"captured_at": frame.captured_at,
|
| 229 |
-
"attribution": task.attribution,
|
| 230 |
-
"n_frames": len(task.frames),
|
| 231 |
-
"start_frame": task.start_frame,
|
| 232 |
-
# Per-frame headings only. Coordinates are withheld for frames
|
| 233 |
-
# other than the start, which is the one the guess is scored
|
| 234 |
-
# against and therefore already revealed to a human player.
|
| 235 |
-
"frames": [
|
| 236 |
-
{
|
| 237 |
-
"index": i,
|
| 238 |
-
"compass_angle": f.compass_angle,
|
| 239 |
-
"captured_at": f.captured_at,
|
| 240 |
-
}
|
| 241 |
-
for i, f in enumerate(task.frames)
|
| 242 |
-
],
|
| 243 |
-
}
|
| 244 |
-
)
|
| 245 |
-
|
| 246 |
-
def _pano_response(task_index: int, frame_index: int | None, split: str | None):
|
| 247 |
-
"""Resolve one frame's panorama file, fetching it if necessary."""
|
| 248 |
-
resolved = _backend(split)
|
| 249 |
-
try:
|
| 250 |
-
task = resolved.task(task_index)
|
| 251 |
-
except IndexError as exc:
|
| 252 |
-
raise HTTPException(status_code=404, detail=str(exc)) from exc
|
| 253 |
-
position = task.start_frame if frame_index is None else frame_index
|
| 254 |
-
if not 0 <= position < len(task.frames):
|
| 255 |
-
raise HTTPException(
|
| 256 |
-
status_code=404,
|
| 257 |
-
detail=(
|
| 258 |
-
f"frame {position} out of range for task {task_index} "
|
| 259 |
-
f"with {len(task.frames)} frames"
|
| 260 |
-
),
|
| 261 |
-
)
|
| 262 |
-
frame = task.frames[position]
|
| 263 |
-
try:
|
| 264 |
-
resolved.load_pano(frame.image_id)
|
| 265 |
-
except Exception as exc:
|
| 266 |
-
raise HTTPException(status_code=503, detail=str(exc)) from exc
|
| 267 |
-
return FileResponse(
|
| 268 |
-
pathlib.Path(CACHE_DIR) / f"{frame.image_id}.jpg",
|
| 269 |
-
media_type="image/jpeg",
|
| 270 |
-
)
|
| 271 |
-
|
| 272 |
-
@application.get(
|
| 273 |
-
"/geoguesser/pano/{task_index}",
|
| 274 |
-
tags=["geoguesser"],
|
| 275 |
-
response_class=FileResponse,
|
| 276 |
-
)
|
| 277 |
-
async def geoguesser_pano(task_index: int, split: str | None = Query(None)):
|
| 278 |
-
"""The task's starting panorama, for the browser viewer."""
|
| 279 |
-
return _pano_response(task_index, None, split)
|
| 280 |
-
|
| 281 |
-
@application.get(
|
| 282 |
-
"/geoguesser/pano/{task_index}/{frame_index}",
|
| 283 |
-
tags=["geoguesser"],
|
| 284 |
-
response_class=FileResponse,
|
| 285 |
-
)
|
| 286 |
-
async def geoguesser_pano_frame(
|
| 287 |
-
task_index: int, frame_index: int, split: str | None = Query(None)
|
| 288 |
-
):
|
| 289 |
-
"""One specific frame's panorama, so the viewer can follow `move()`."""
|
| 290 |
-
return _pano_response(task_index, frame_index, split)
|
| 291 |
-
|
| 292 |
-
@application.get(
|
| 293 |
-
"/geoguesser/play", tags=["geoguesser"], response_class=HTMLResponse
|
| 294 |
-
)
|
| 295 |
-
async def geoguesser_play(split: str | None = Query(None)):
|
| 296 |
-
"""The standalone play page, also embedded in the Gradio tab."""
|
| 297 |
-
from .gradio_ui import play_page_html
|
| 298 |
-
|
| 299 |
-
return HTMLResponse(play_page_html(catalog.list_splits(), split))
|
| 300 |
-
|
| 301 |
-
@application.get(
|
| 302 |
-
"/geoguesser/tasks", tags=["geoguesser"], response_class=JSONResponse
|
| 303 |
-
)
|
| 304 |
-
async def geoguesser_tasks():
|
| 305 |
-
"""Which splits exist and how many tasks each holds."""
|
| 306 |
-
splits = catalog.list_splits()
|
| 307 |
-
return JSONResponse(
|
| 308 |
-
{
|
| 309 |
-
"splits": splits,
|
| 310 |
-
"default_split": ACTIVE_DEFAULT_SPLIT,
|
| 311 |
-
# Kept so an older play page still finds a count.
|
| 312 |
-
"n_tasks": next(
|
| 313 |
-
s["num_tasks"] for s in splits if s["name"] == ACTIVE_DEFAULT_SPLIT
|
| 314 |
-
),
|
| 315 |
-
}
|
| 316 |
-
)
|
| 317 |
-
|
| 318 |
-
|
| 319 |
-
app = _build_app()
|
| 320 |
-
|
| 321 |
-
try:
|
| 322 |
-
_attach_play_routes(app)
|
| 323 |
-
except Exception as exc: # pragma: no cover - index may be absent in CI
|
| 324 |
-
logger.warning("play routes unavailable: %r", exc)
|
| 325 |
-
|
| 326 |
-
|
| 327 |
-
def main(host: str = "0.0.0.0", port: int = 8000) -> None:
|
| 328 |
-
"""
|
| 329 |
-
Entry point for running the server without Docker.
|
| 330 |
-
|
| 331 |
-
Args:
|
| 332 |
-
host (`str`, *optional*, defaults to `"0.0.0.0"`):
|
| 333 |
-
Address to bind.
|
| 334 |
-
port (`int`, *optional*, defaults to `8000`):
|
| 335 |
-
Port to listen on.
|
| 336 |
-
"""
|
| 337 |
-
import uvicorn
|
| 338 |
-
|
| 339 |
-
uvicorn.run(app, host=host, port=port)
|
| 340 |
-
|
| 341 |
-
|
| 342 |
-
if __name__ == "__main__":
|
| 343 |
-
import argparse
|
| 344 |
-
|
| 345 |
-
parser = argparse.ArgumentParser()
|
| 346 |
-
parser.add_argument("--port", type=int, default=8000)
|
| 347 |
-
parser.add_argument("--host", default="0.0.0.0")
|
| 348 |
-
args = parser.parse_args()
|
| 349 |
-
main(host=args.host, port=args.port)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/backends/__init__.py
DELETED
|
@@ -1,7 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Imagery backends for the GeoGuesser environment."""
|
| 4 |
-
|
| 5 |
-
from .base import Frame, PanoramaBackend, Task
|
| 6 |
-
|
| 7 |
-
__all__ = ["Frame", "PanoramaBackend", "Task"]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/backends/base.py
DELETED
|
@@ -1,127 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""The contract every imagery backend implements.
|
| 4 |
-
|
| 5 |
-
Backends differ only in where pixels come from and whether the location can be
|
| 6 |
-
walked. Everything above them — scoring, parsing, the pin loop, the map — is
|
| 7 |
-
shared, so a policy trained against one backend runs unmodified against
|
| 8 |
-
another.
|
| 9 |
-
|
| 10 |
-
A backend that cannot do something reports it through `supports_look` or
|
| 11 |
-
`supports_move`; the environment then simply does not register the
|
| 12 |
-
corresponding tools. Registering a tool that always fails would only teach a
|
| 13 |
-
policy to spend its step budget discovering that.
|
| 14 |
-
"""
|
| 15 |
-
|
| 16 |
-
from __future__ import annotations
|
| 17 |
-
|
| 18 |
-
from dataclasses import dataclass, field
|
| 19 |
-
from typing import Any, Protocol, runtime_checkable
|
| 20 |
-
|
| 21 |
-
from PIL import Image
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
@dataclass(frozen=True)
|
| 25 |
-
class Frame:
|
| 26 |
-
"""One captured position within a sequence.
|
| 27 |
-
|
| 28 |
-
Attributes:
|
| 29 |
-
image_id (`str`):
|
| 30 |
-
Provider-side identifier, also the cache filename.
|
| 31 |
-
lat (`float`):
|
| 32 |
-
Latitude in degrees.
|
| 33 |
-
lon (`float`):
|
| 34 |
-
Longitude in degrees.
|
| 35 |
-
compass_angle (`float`):
|
| 36 |
-
Heading the camera faced, in degrees clockwise from true north.
|
| 37 |
-
captured_at (`str`):
|
| 38 |
-
Capture month as `YYYY-MM`.
|
| 39 |
-
is_pano (`bool`, *optional*, defaults to `True`):
|
| 40 |
-
Whether this frame is a full 360-degree panorama.
|
| 41 |
-
"""
|
| 42 |
-
|
| 43 |
-
image_id: str
|
| 44 |
-
lat: float
|
| 45 |
-
lon: float
|
| 46 |
-
compass_angle: float
|
| 47 |
-
captured_at: str
|
| 48 |
-
is_pano: bool = True
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
@dataclass(frozen=True)
|
| 52 |
-
class Task:
|
| 53 |
-
"""One episode's location, frozen at index build time.
|
| 54 |
-
|
| 55 |
-
Attributes:
|
| 56 |
-
task_index (`int`):
|
| 57 |
-
Position in the task list. Stable, and what `reset` selects on.
|
| 58 |
-
task_id (`str`):
|
| 59 |
-
Human-readable identifier, unique within the index.
|
| 60 |
-
frames (`list[Frame]`):
|
| 61 |
-
Ordered frames of the captured sequence, walkable with `move`.
|
| 62 |
-
start_frame (`int`):
|
| 63 |
-
Index into `frames` where the episode begins.
|
| 64 |
-
country (`str`):
|
| 65 |
-
ISO-3166 alpha-2 code of the true location. Ground truth — never
|
| 66 |
-
placed in an observation before the guess.
|
| 67 |
-
sequence_id (`str`):
|
| 68 |
-
Provider-side sequence identifier.
|
| 69 |
-
provider (`str`, *optional*, defaults to `"mapillary"`):
|
| 70 |
-
Which backend can resolve this task's imagery.
|
| 71 |
-
attribution (`dict`, *optional*):
|
| 72 |
-
Creator credit, required by the CC-BY-SA licence on the imagery.
|
| 73 |
-
meta (`dict`, *optional*):
|
| 74 |
-
Anything else the builder recorded — camera make and model,
|
| 75 |
-
quality score, checksums.
|
| 76 |
-
"""
|
| 77 |
-
|
| 78 |
-
task_index: int
|
| 79 |
-
task_id: str
|
| 80 |
-
frames: list[Frame]
|
| 81 |
-
start_frame: int
|
| 82 |
-
country: str
|
| 83 |
-
sequence_id: str
|
| 84 |
-
provider: str = "mapillary"
|
| 85 |
-
attribution: dict[str, Any] = field(default_factory=dict)
|
| 86 |
-
meta: dict[str, Any] = field(default_factory=dict)
|
| 87 |
-
|
| 88 |
-
@property
|
| 89 |
-
def truth(self) -> tuple[float, float]:
|
| 90 |
-
"""Ground-truth `(lat, lon)` of the starting frame."""
|
| 91 |
-
frame = self.frames[self.start_frame]
|
| 92 |
-
return frame.lat, frame.lon
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
@runtime_checkable
|
| 96 |
-
class PanoramaBackend(Protocol):
|
| 97 |
-
"""Resolves tasks to imagery and answers what the agent may do."""
|
| 98 |
-
|
| 99 |
-
@property
|
| 100 |
-
def n_tasks(self) -> int:
|
| 101 |
-
"""Number of tasks in the frozen index."""
|
| 102 |
-
...
|
| 103 |
-
|
| 104 |
-
@property
|
| 105 |
-
def supports_look(self) -> bool:
|
| 106 |
-
"""Whether views can be rendered at arbitrary headings."""
|
| 107 |
-
...
|
| 108 |
-
|
| 109 |
-
@property
|
| 110 |
-
def supports_move(self) -> bool:
|
| 111 |
-
"""Whether the location can be walked along a sequence."""
|
| 112 |
-
...
|
| 113 |
-
|
| 114 |
-
def task(self, task_index: int) -> Task:
|
| 115 |
-
"""Return the task at `task_index`."""
|
| 116 |
-
...
|
| 117 |
-
|
| 118 |
-
def render_view(
|
| 119 |
-
self,
|
| 120 |
-
task: Task,
|
| 121 |
-
frame_index: int,
|
| 122 |
-
heading_deg: float,
|
| 123 |
-
pitch_deg: float,
|
| 124 |
-
fov_deg: float,
|
| 125 |
-
) -> Image.Image:
|
| 126 |
-
"""Render what the camera sees from one frame."""
|
| 127 |
-
...
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/backends/panorama.py
DELETED
|
@@ -1,399 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Mapillary-backed panoramas, read from a frozen index and a local cache.
|
| 4 |
-
|
| 5 |
-
The index (`tasks/pano_v1.jsonl`) is self-contained: it carries every frame's
|
| 6 |
-
coordinates, heading and capture date, so the movement graph resolves with no
|
| 7 |
-
network access at all. Only image *bytes* may need fetching, and only on a
|
| 8 |
-
cache miss, because Mapillary's `thumb_*_url` values are expiring signed CDN
|
| 9 |
-
URLs and cannot be stored in the index.
|
| 10 |
-
|
| 11 |
-
Once a task's frames are cached, episodes are byte-identical on repeat — the
|
| 12 |
-
property a GRPO group depends on.
|
| 13 |
-
"""
|
| 14 |
-
|
| 15 |
-
from __future__ import annotations
|
| 16 |
-
|
| 17 |
-
import hashlib
|
| 18 |
-
import json
|
| 19 |
-
import logging
|
| 20 |
-
import os
|
| 21 |
-
import pathlib
|
| 22 |
-
import urllib.parse
|
| 23 |
-
import urllib.request
|
| 24 |
-
|
| 25 |
-
from PIL import Image
|
| 26 |
-
|
| 27 |
-
from ..render.pano import look
|
| 28 |
-
from .base import Frame, Task
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
logger = logging.getLogger(__name__)
|
| 32 |
-
|
| 33 |
-
GRAPH_API = "https://graph.mapillary.com"
|
| 34 |
-
FETCH_TIMEOUT_S = 60.0
|
| 35 |
-
|
| 36 |
-
# Zooming into a 2048x1024 panorama is resolution-starved: a 30-degree view
|
| 37 |
-
# samples only ~170 source pixels, so narrowing the field of view barely adds
|
| 38 |
-
# detail (measured mean gradient 6.60 at 90 degrees against 7.03 at 30). The
|
| 39 |
-
# 7680x3840 original roughly doubles it instead (10.23 against 14.87), which is
|
| 40 |
-
# what makes reading a distant sign possible at all.
|
| 41 |
-
#
|
| 42 |
-
# So two derivatives are cached per panorama: the 2048 for wide views, which
|
| 43 |
-
# renders in ~30 ms, and the original for zoomed views at ~70 ms and 3.3 MB.
|
| 44 |
-
THUMB_VARIANT = "thumb_2048_url"
|
| 45 |
-
ORIGINAL_VARIANT = "thumb_original_url"
|
| 46 |
-
HIRES_FOV_DEG = 45.0
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
class MissingImageError(RuntimeError):
|
| 50 |
-
"""A frame is absent from the cache and cannot be fetched."""
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
_INDEX_CACHE: dict[tuple[str, float, int], list[Task]] = {}
|
| 54 |
-
"""Parsed indexes keyed by path, mtime and size, so edits invalidate the entry."""
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
def load_index(index_path: str | pathlib.Path) -> list[Task]:
|
| 58 |
-
"""
|
| 59 |
-
Parse a task index, reusing an already-parsed copy when possible.
|
| 60 |
-
|
| 61 |
-
The Task API routes construct a throwaway environment per request and a
|
| 62 |
-
5,000-task index is roughly 31 MB, so re-parsing it on every call is not
|
| 63 |
-
affordable. Keying on mtime and size means rebuilding an index in place
|
| 64 |
-
invalidates the entry instead of serving stale tasks.
|
| 65 |
-
|
| 66 |
-
Args:
|
| 67 |
-
index_path (`str` or `pathlib.Path`):
|
| 68 |
-
JSONL index written by `scripts/build_tasks.py`.
|
| 69 |
-
|
| 70 |
-
Returns:
|
| 71 |
-
`list[Task]`: Tasks ordered by `task_index`. The list is shared between
|
| 72 |
-
callers, so treat it as read-only.
|
| 73 |
-
|
| 74 |
-
Raises:
|
| 75 |
-
FileNotFoundError: When the index does not exist.
|
| 76 |
-
ValueError: When the index is empty, or its task indices are not
|
| 77 |
-
contiguous from zero.
|
| 78 |
-
"""
|
| 79 |
-
path = pathlib.Path(index_path)
|
| 80 |
-
if not path.exists():
|
| 81 |
-
raise FileNotFoundError(
|
| 82 |
-
f"Task index not found: {path}. Build one with scripts/build_tasks.py."
|
| 83 |
-
)
|
| 84 |
-
stat = path.stat()
|
| 85 |
-
key = (str(path.resolve()), stat.st_mtime, stat.st_size)
|
| 86 |
-
cached = _INDEX_CACHE.get(key)
|
| 87 |
-
if cached is not None:
|
| 88 |
-
return cached
|
| 89 |
-
|
| 90 |
-
tasks: list[Task] = []
|
| 91 |
-
for line in path.read_text().splitlines():
|
| 92 |
-
line = line.strip()
|
| 93 |
-
if not line:
|
| 94 |
-
continue
|
| 95 |
-
row = json.loads(line)
|
| 96 |
-
frames = [
|
| 97 |
-
Frame(
|
| 98 |
-
image_id=str(f["image_id"]),
|
| 99 |
-
lat=float(f["lat"]),
|
| 100 |
-
lon=float(f["lon"]),
|
| 101 |
-
compass_angle=float(f.get("compass_angle", 0.0)),
|
| 102 |
-
captured_at=str(f.get("captured_at", "")),
|
| 103 |
-
is_pano=bool(f.get("is_pano", True)),
|
| 104 |
-
)
|
| 105 |
-
for f in row["frames"]
|
| 106 |
-
]
|
| 107 |
-
tasks.append(
|
| 108 |
-
Task(
|
| 109 |
-
task_index=int(row["task_index"]),
|
| 110 |
-
task_id=str(row["task_id"]),
|
| 111 |
-
frames=frames,
|
| 112 |
-
start_frame=int(row.get("start_frame", 0)),
|
| 113 |
-
country=str(row.get("country", "")),
|
| 114 |
-
sequence_id=str(row.get("sequence_id", "")),
|
| 115 |
-
provider=str(row.get("provider", "mapillary")),
|
| 116 |
-
attribution=row.get("attribution", {}),
|
| 117 |
-
meta=row.get("meta", {}),
|
| 118 |
-
)
|
| 119 |
-
)
|
| 120 |
-
if not tasks:
|
| 121 |
-
raise ValueError(f"Task index {path} is empty.")
|
| 122 |
-
tasks.sort(key=lambda t: t.task_index)
|
| 123 |
-
for position, task in enumerate(tasks):
|
| 124 |
-
if task.task_index != position:
|
| 125 |
-
raise ValueError(
|
| 126 |
-
"Task indices must be contiguous from 0; found "
|
| 127 |
-
f"{task.task_index} at position {position}."
|
| 128 |
-
)
|
| 129 |
-
_INDEX_CACHE[key] = tasks
|
| 130 |
-
return tasks
|
| 131 |
-
|
| 132 |
-
|
| 133 |
-
class PanoramaBackend:
|
| 134 |
-
"""Serve panoramas from a frozen task index plus a disk cache.
|
| 135 |
-
|
| 136 |
-
Args:
|
| 137 |
-
index_path (`str` or `pathlib.Path`):
|
| 138 |
-
JSONL task index produced by `scripts/build_pano_tasks.py`.
|
| 139 |
-
cache_dir (`str` or `pathlib.Path`):
|
| 140 |
-
Directory holding cached JPEGs, named `<image_id>.jpg`.
|
| 141 |
-
access_token (`str`, *optional*):
|
| 142 |
-
Mapillary token, used only to fill cache misses. When absent, a
|
| 143 |
-
miss raises [`MissingImageError`] instead of reaching the network.
|
| 144 |
-
allow_fetch (`bool`, *optional*, defaults to `True`):
|
| 145 |
-
Set `False` to guarantee an episode never touches the network.
|
| 146 |
-
hires_zoom (`bool`, *optional*, defaults to `True`):
|
| 147 |
-
Render fields of view at or below `hires_fov_deg` from the
|
| 148 |
-
full-resolution original, so zooming actually resolves detail.
|
| 149 |
-
Falls back to the 2048 derivative when no original exists.
|
| 150 |
-
hires_fov_deg (`float`, *optional*, defaults to `45.0`):
|
| 151 |
-
Field of view at or below which the original is used.
|
| 152 |
-
verify_checksums (`bool`, *optional*, defaults to `False`):
|
| 153 |
-
Verify each cached start frame against the sha256 recorded at build
|
| 154 |
-
time. Used by frozen evals to detect drift.
|
| 155 |
-
|
| 156 |
-
Examples:
|
| 157 |
-
|
| 158 |
-
```python
|
| 159 |
-
backend = PanoramaBackend("tasks/pano_v1.jsonl", "data/panos")
|
| 160 |
-
task = backend.task(0)
|
| 161 |
-
view = backend.render_view(task, task.start_frame, 90.0, 0.0, 90.0)
|
| 162 |
-
```
|
| 163 |
-
"""
|
| 164 |
-
|
| 165 |
-
def __init__(
|
| 166 |
-
self,
|
| 167 |
-
index_path: str | pathlib.Path,
|
| 168 |
-
cache_dir: str | pathlib.Path,
|
| 169 |
-
access_token: str | None = None,
|
| 170 |
-
allow_fetch: bool = True,
|
| 171 |
-
verify_checksums: bool = False,
|
| 172 |
-
hires_zoom: bool = True,
|
| 173 |
-
hires_fov_deg: float = HIRES_FOV_DEG,
|
| 174 |
-
):
|
| 175 |
-
self._index_path = pathlib.Path(index_path)
|
| 176 |
-
self._cache_dir = pathlib.Path(cache_dir)
|
| 177 |
-
self._cache_dir.mkdir(parents=True, exist_ok=True)
|
| 178 |
-
self._token = access_token or os.environ.get("MAPILLARY_API_KEY")
|
| 179 |
-
self._allow_fetch = allow_fetch
|
| 180 |
-
self._verify_checksums = verify_checksums
|
| 181 |
-
self._hires_zoom = hires_zoom
|
| 182 |
-
self._hires_fov_deg = hires_fov_deg
|
| 183 |
-
self._tasks = self._load_index()
|
| 184 |
-
logger.info(
|
| 185 |
-
"loaded %d tasks from %s (cache: %s, fetch: %s)",
|
| 186 |
-
len(self._tasks),
|
| 187 |
-
self._index_path,
|
| 188 |
-
self._cache_dir,
|
| 189 |
-
"on" if self._allow_fetch and self._token else "off",
|
| 190 |
-
)
|
| 191 |
-
|
| 192 |
-
# -- index ------------------------------------------------------------
|
| 193 |
-
|
| 194 |
-
def _load_index(self) -> list[Task]:
|
| 195 |
-
return load_index(self._index_path)
|
| 196 |
-
|
| 197 |
-
# -- capabilities -----------------------------------------------------
|
| 198 |
-
|
| 199 |
-
@property
|
| 200 |
-
def n_tasks(self) -> int:
|
| 201 |
-
"""Number of tasks in the frozen index."""
|
| 202 |
-
return len(self._tasks)
|
| 203 |
-
|
| 204 |
-
@property
|
| 205 |
-
def supports_look(self) -> bool:
|
| 206 |
-
"""True — every task in this index is a 360-degree panorama."""
|
| 207 |
-
return True
|
| 208 |
-
|
| 209 |
-
@property
|
| 210 |
-
def supports_move(self) -> bool:
|
| 211 |
-
"""Whether any task has more than one frame to walk between."""
|
| 212 |
-
return any(len(t.frames) > 1 for t in self._tasks)
|
| 213 |
-
|
| 214 |
-
def task(self, task_index: int) -> Task:
|
| 215 |
-
"""
|
| 216 |
-
Return the task at `task_index`.
|
| 217 |
-
|
| 218 |
-
Args:
|
| 219 |
-
task_index (`int`):
|
| 220 |
-
Position in the frozen index.
|
| 221 |
-
|
| 222 |
-
Returns:
|
| 223 |
-
[`Task`]: The task, including its full frame list.
|
| 224 |
-
"""
|
| 225 |
-
if not 0 <= task_index < len(self._tasks):
|
| 226 |
-
raise IndexError(
|
| 227 |
-
f"task_index {task_index} out of range for {len(self._tasks)} tasks."
|
| 228 |
-
)
|
| 229 |
-
return self._tasks[task_index]
|
| 230 |
-
|
| 231 |
-
# -- imagery ----------------------------------------------------------
|
| 232 |
-
|
| 233 |
-
def _cache_path(self, image_id: str, hires: bool = False) -> pathlib.Path:
|
| 234 |
-
suffix = ".orig.jpg" if hires else ".jpg"
|
| 235 |
-
return self._cache_dir / f"{image_id}{suffix}"
|
| 236 |
-
|
| 237 |
-
def _fetch(self, image_id: str, hires: bool = False) -> bytes:
|
| 238 |
-
if not self._allow_fetch:
|
| 239 |
-
raise MissingImageError(
|
| 240 |
-
f"Image {image_id} is not cached and fetching is disabled. "
|
| 241 |
-
"Warm the cache with scripts/build_pano_tasks.py --prefetch-frames."
|
| 242 |
-
)
|
| 243 |
-
if not self._token:
|
| 244 |
-
raise MissingImageError(
|
| 245 |
-
f"Image {image_id} is not cached and MAPILLARY_API_KEY is unset, "
|
| 246 |
-
"so it cannot be fetched."
|
| 247 |
-
)
|
| 248 |
-
variant = ORIGINAL_VARIANT if hires else THUMB_VARIANT
|
| 249 |
-
meta_url = f"{GRAPH_API}/{image_id}?" + urllib.parse.urlencode(
|
| 250 |
-
{"fields": variant, "access_token": self._token}
|
| 251 |
-
)
|
| 252 |
-
with urllib.request.urlopen(meta_url, timeout=FETCH_TIMEOUT_S) as response:
|
| 253 |
-
thumb_url = json.loads(response.read()).get(variant)
|
| 254 |
-
if not thumb_url:
|
| 255 |
-
raise MissingImageError(
|
| 256 |
-
f"Mapillary returned no {variant} for {image_id}; its "
|
| 257 |
-
"derivatives may have been removed."
|
| 258 |
-
)
|
| 259 |
-
with urllib.request.urlopen(thumb_url, timeout=FETCH_TIMEOUT_S) as response:
|
| 260 |
-
return response.read()
|
| 261 |
-
|
| 262 |
-
def load_pano(
|
| 263 |
-
self,
|
| 264 |
-
image_id: str,
|
| 265 |
-
expected_sha256: str | None = None,
|
| 266 |
-
hires: bool = False,
|
| 267 |
-
) -> Image.Image:
|
| 268 |
-
"""
|
| 269 |
-
Return a panorama, fetching and caching it if necessary.
|
| 270 |
-
|
| 271 |
-
Args:
|
| 272 |
-
image_id (`str`):
|
| 273 |
-
Provider-side image identifier.
|
| 274 |
-
expected_sha256 (`str`, *optional*):
|
| 275 |
-
Checksum recorded at build time. Verified only when the backend
|
| 276 |
-
was constructed with `verify_checksums=True`, and only for
|
| 277 |
-
the 2048 derivative, which is what the index records.
|
| 278 |
-
hires (`bool`, *optional*, defaults to `False`):
|
| 279 |
-
Load the full-resolution original instead of the 2048
|
| 280 |
-
derivative.
|
| 281 |
-
|
| 282 |
-
Returns:
|
| 283 |
-
`PIL.Image.Image`: The equirectangular panorama.
|
| 284 |
-
"""
|
| 285 |
-
path = self._cache_path(image_id, hires=hires)
|
| 286 |
-
if not path.exists():
|
| 287 |
-
payload = self._fetch(image_id, hires=hires)
|
| 288 |
-
path.write_bytes(payload)
|
| 289 |
-
logger.info(
|
| 290 |
-
"cached %s%s (%.0f KB)",
|
| 291 |
-
image_id,
|
| 292 |
-
" at full resolution" if hires else "",
|
| 293 |
-
len(payload) / 1024,
|
| 294 |
-
)
|
| 295 |
-
if self._verify_checksums and expected_sha256 and not hires:
|
| 296 |
-
actual = hashlib.sha256(path.read_bytes()).hexdigest()
|
| 297 |
-
if actual != expected_sha256:
|
| 298 |
-
raise MissingImageError(
|
| 299 |
-
f"Checksum mismatch for {image_id}: index recorded "
|
| 300 |
-
f"{expected_sha256[:12]}, cache holds {actual[:12]}. The "
|
| 301 |
-
"upstream image changed; this task is no longer comparable."
|
| 302 |
-
)
|
| 303 |
-
return Image.open(path)
|
| 304 |
-
|
| 305 |
-
def render_view(
|
| 306 |
-
self,
|
| 307 |
-
task: Task,
|
| 308 |
-
frame_index: int,
|
| 309 |
-
heading_deg: float,
|
| 310 |
-
pitch_deg: float = 0.0,
|
| 311 |
-
fov_deg: float = 90.0,
|
| 312 |
-
) -> Image.Image:
|
| 313 |
-
"""
|
| 314 |
-
Render what the camera sees from one frame of a task.
|
| 315 |
-
|
| 316 |
-
Headings are absolute: `0` is true north, obtained by offsetting the
|
| 317 |
-
request by the frame's own `compass_angle`. That keeps `look(0)`
|
| 318 |
-
meaning the same thing in every task.
|
| 319 |
-
|
| 320 |
-
Args:
|
| 321 |
-
task ([`Task`]):
|
| 322 |
-
The task being played.
|
| 323 |
-
frame_index (`int`):
|
| 324 |
-
Which frame of the sequence the agent stands on.
|
| 325 |
-
heading_deg (`float`):
|
| 326 |
-
Absolute compass heading in degrees.
|
| 327 |
-
pitch_deg (`float`, *optional*, defaults to `0.0`):
|
| 328 |
-
Vertical angle in degrees.
|
| 329 |
-
fov_deg (`float`, *optional*, defaults to `90.0`):
|
| 330 |
-
Horizontal field of view in degrees.
|
| 331 |
-
|
| 332 |
-
Returns:
|
| 333 |
-
`PIL.Image.Image`: The rendered view.
|
| 334 |
-
"""
|
| 335 |
-
frame = task.frames[frame_index]
|
| 336 |
-
checksums = task.meta.get("sha256", {})
|
| 337 |
-
want_hires = self._hires_zoom and fov_deg <= self._hires_fov_deg
|
| 338 |
-
try:
|
| 339 |
-
pano = self.load_pano(
|
| 340 |
-
frame.image_id, checksums.get(frame.image_id), hires=want_hires
|
| 341 |
-
)
|
| 342 |
-
except MissingImageError:
|
| 343 |
-
if not want_hires:
|
| 344 |
-
raise
|
| 345 |
-
# A missing original must not end an episode; a soft view beats a
|
| 346 |
-
# failed step.
|
| 347 |
-
logger.warning(
|
| 348 |
-
"no full-resolution original for %s; zooming on the 2048 "
|
| 349 |
-
"derivative instead",
|
| 350 |
-
frame.image_id,
|
| 351 |
-
)
|
| 352 |
-
pano = self.load_pano(frame.image_id, checksums.get(frame.image_id))
|
| 353 |
-
return look(pano, heading_deg + frame.compass_angle, pitch_deg, fov_deg)
|
| 354 |
-
|
| 355 |
-
# -- navigation -------------------------------------------------------
|
| 356 |
-
|
| 357 |
-
def step_along(
|
| 358 |
-
self, task: Task, frame_index: int, direction: str, meters: float
|
| 359 |
-
) -> tuple[int, float]:
|
| 360 |
-
"""
|
| 361 |
-
Walk the sequence and report where the agent actually ended up.
|
| 362 |
-
|
| 363 |
-
Frame spacing is irregular — measured around 3.3 m on Mapillary
|
| 364 |
-
sequences — so the requested distance is consumed frame by frame and
|
| 365 |
-
the realised distance is returned rather than assumed.
|
| 366 |
-
|
| 367 |
-
Args:
|
| 368 |
-
task ([`Task`]):
|
| 369 |
-
The task being played.
|
| 370 |
-
frame_index (`int`):
|
| 371 |
-
Current position in `task.frames`.
|
| 372 |
-
direction (`str`):
|
| 373 |
-
`"forward"` or `"backward"`.
|
| 374 |
-
meters (`float`):
|
| 375 |
-
Requested distance in metres.
|
| 376 |
-
|
| 377 |
-
Returns:
|
| 378 |
-
`tuple[int, float]` with:
|
| 379 |
-
- the new frame index, unchanged at a dead end
|
| 380 |
-
- metres actually travelled
|
| 381 |
-
"""
|
| 382 |
-
from ..scoring import haversine_km
|
| 383 |
-
|
| 384 |
-
step = 1 if direction == "forward" else -1
|
| 385 |
-
current = frame_index
|
| 386 |
-
travelled = 0.0
|
| 387 |
-
while travelled < meters:
|
| 388 |
-
nxt = current + step
|
| 389 |
-
if not 0 <= nxt < len(task.frames):
|
| 390 |
-
break
|
| 391 |
-
a, b = task.frames[current], task.frames[nxt]
|
| 392 |
-
travelled += haversine_km(a.lat, a.lon, b.lat, b.lon) * 1000.0
|
| 393 |
-
current = nxt
|
| 394 |
-
return current, travelled
|
| 395 |
-
|
| 396 |
-
def can_move(self, task: Task, frame_index: int, direction: str) -> bool:
|
| 397 |
-
"""Whether a frame exists in `direction` from the current position."""
|
| 398 |
-
step = 1 if direction == "forward" else -1
|
| 399 |
-
return 0 <= frame_index + step < len(task.frames)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/geoguesser_environment.py
DELETED
|
@@ -1,998 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""The GeoGuesser environment.
|
| 4 |
-
|
| 5 |
-
Exposes its tools over MCP so an agent can look around, walk, and check
|
| 6 |
-
candidate coordinates on a map before committing to a guess. Non-MCP
|
| 7 |
-
structured actions route through `_step_impl`, which is what makes single-shot
|
| 8 |
-
GRPO and the agentic loop the same environment rather than two.
|
| 9 |
-
|
| 10 |
-
Two rules are load-bearing:
|
| 11 |
-
|
| 12 |
-
- Pin feedback describes only where the agent pointed. Any signal about the
|
| 13 |
-
target would make binary search optimal, and the benchmark would measure
|
| 14 |
-
bisection instead of geography.
|
| 15 |
-
- Tools the backend cannot serve are never registered, rather than registered
|
| 16 |
-
and failing, so a policy does not learn to spend steps on dead ends.
|
| 17 |
-
"""
|
| 18 |
-
|
| 19 |
-
from __future__ import annotations
|
| 20 |
-
|
| 21 |
-
import logging
|
| 22 |
-
import random
|
| 23 |
-
import uuid
|
| 24 |
-
from typing import Any
|
| 25 |
-
|
| 26 |
-
from fastmcp import FastMCP
|
| 27 |
-
from openenv.core.env_server.mcp_environment import MCPEnvironment
|
| 28 |
-
from openenv.core.env_server.types import Action, Observation
|
| 29 |
-
|
| 30 |
-
from ..models import (
|
| 31 |
-
EpisodeMode,
|
| 32 |
-
from_wire,
|
| 33 |
-
GeoGuesserAction,
|
| 34 |
-
GeoGuesserObservation,
|
| 35 |
-
GeoGuesserState,
|
| 36 |
-
GuessAction,
|
| 37 |
-
LookAction,
|
| 38 |
-
MeasureAction,
|
| 39 |
-
MoveAction,
|
| 40 |
-
PanAction,
|
| 41 |
-
Pin,
|
| 42 |
-
PinAction,
|
| 43 |
-
RewardMode,
|
| 44 |
-
TypedAction,
|
| 45 |
-
ViewMapAction,
|
| 46 |
-
ZoomAction,
|
| 47 |
-
)
|
| 48 |
-
from .backends.base import Task
|
| 49 |
-
from .backends.panorama import PanoramaBackend
|
| 50 |
-
from .parser import parse_guess
|
| 51 |
-
from .render.minimap import (
|
| 52 |
-
describe_pin,
|
| 53 |
-
locate,
|
| 54 |
-
render_map,
|
| 55 |
-
street_detail_enabled,
|
| 56 |
-
street_fetch_failed,
|
| 57 |
-
)
|
| 58 |
-
from .render.pano import to_base64
|
| 59 |
-
from .scoring import action_cost, compute_reward, haversine_km, verdict
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
logger = logging.getLogger(__name__)
|
| 63 |
-
|
| 64 |
-
PROMPT = (
|
| 65 |
-
"You are dropped at an unknown street-level location somewhere in the "
|
| 66 |
-
"world. Work out where you are.\n\n"
|
| 67 |
-
"Available tools: {tools}\n\n"
|
| 68 |
-
"Looking around and checking the map cost a little reward each, so gather "
|
| 69 |
-
"the evidence you need and then commit. You have {steps} actions. "
|
| 70 |
-
"Placing a pin shows you where on the map that coordinate falls - it "
|
| 71 |
-
"tells you nothing about whether you are right. Finish with "
|
| 72 |
-
"submit_guess."
|
| 73 |
-
)
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
class UnknownSplitError(KeyError, IndexError):
|
| 77 |
-
"""A split name that is not configured.
|
| 78 |
-
|
| 79 |
-
Inherits both exception types deliberately. `KeyError` is what a Python
|
| 80 |
-
caller expects from a bad name, while the core Task API dispatcher maps
|
| 81 |
-
only `NotImplementedError` and `IndexError` onto HTTP status codes -- so
|
| 82 |
-
without `IndexError` an unknown split surfaces as a 500 instead of a 400.
|
| 83 |
-
"""
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
class GeoGuesserEnvironment(MCPEnvironment):
|
| 87 |
-
"""A GeoGuessr-style geolocation episode.
|
| 88 |
-
|
| 89 |
-
Args:
|
| 90 |
-
index_path (`str`):
|
| 91 |
-
JSONL task index built by `scripts/build_pano_tasks.py`.
|
| 92 |
-
cache_dir (`str`):
|
| 93 |
-
Directory of cached panorama JPEGs.
|
| 94 |
-
episode_mode (`str`, *optional*, defaults to `"agentic"`):
|
| 95 |
-
One of `"agentic"`, `"single_shot"` or `"nmpz"`.
|
| 96 |
-
max_steps (`int`, *optional*, defaults to `24`):
|
| 97 |
-
Actions allowed before the episode is cut off. Twelve is tight
|
| 98 |
-
for an agentic episode: looking in four directions and walking
|
| 99 |
-
a block spends most of it before any reasoning about the map.
|
| 100 |
-
reward_mode (`str`, *optional*, defaults to `"coords"`):
|
| 101 |
-
`"coords"` scores distance; `"country_only"` scores the country.
|
| 102 |
-
max_free_calls (`int`, *optional*, defaults to `24`):
|
| 103 |
-
How many free-tool calls (`measure`) an episode may make before they
|
| 104 |
-
begin consuming the step budget. Free tools return arithmetic on
|
| 105 |
-
coordinates the agent supplied and so reveal nothing, but without a
|
| 106 |
-
cap a policy can issue them indefinitely and never terminate.
|
| 107 |
-
reward_shape (`str`, *optional*, defaults to `"geoguessr"`):
|
| 108 |
-
Distance curve. `"geoguessr"` is the game's own; `"mixture"` adds a
|
| 109 |
-
5000 km scale so a wrong-continent guess still has a gradient. Use
|
| 110 |
-
`"mixture"` for training and `"geoguessr"` for anything you report.
|
| 111 |
-
cost_mode (`str`, *optional*, defaults to `"subtract"`):
|
| 112 |
-
`"subtract"` takes the action cost off the score and floors at zero,
|
| 113 |
-
as the game does. `"multiply"` scales the score by `1 - cost`, which
|
| 114 |
-
preserves the ordering of bad guesses -- required for training, see
|
| 115 |
-
[`~scoring.compute_reward`].
|
| 116 |
-
hide_task_identity (`bool`, *optional*, defaults to `False`):
|
| 117 |
-
Drop `task_index`, `task_id`, `sequence_id` and `attribution` from
|
| 118 |
-
the per-turn observation metadata. **Set this for RL training**: the
|
| 119 |
-
contributor username alone determines the country for 74% of
|
| 120 |
-
training tasks, so leaving it in lets a policy score without looking
|
| 121 |
-
at the image. The terminal observation carries them either way.
|
| 122 |
-
hierarchical_reward (`bool`, *optional*, defaults to `False`):
|
| 123 |
-
Add country and region partial credit to the distance score.
|
| 124 |
-
view_size (`int`, *optional*, defaults to `640`):
|
| 125 |
-
Edge length in pixels of rendered views.
|
| 126 |
-
allow_fetch (`bool`, *optional*, defaults to `True`):
|
| 127 |
-
Whether a cache miss may reach the Mapillary API.
|
| 128 |
-
hires_zoom (`bool`, *optional*, defaults to `True`):
|
| 129 |
-
Render zoomed views from the full-resolution original, so a narrow
|
| 130 |
-
field of view actually resolves detail such as distant signage.
|
| 131 |
-
reveal_map (`bool`, *optional*, defaults to `True`):
|
| 132 |
-
Draw a map of the guess against the truth on the terminal
|
| 133 |
-
observation. Costs about 280 ms, which is the single largest cost in
|
| 134 |
-
an episode, and a training run does not read it — the reward and the
|
| 135 |
-
distance are in the observation either way. Leave it on for evals,
|
| 136 |
-
demos and traces; turn it off for throughput.
|
| 137 |
-
|
| 138 |
-
Examples:
|
| 139 |
-
|
| 140 |
-
```python
|
| 141 |
-
env = GeoGuesserEnvironment("tasks/pano_v1.jsonl", "data/panos")
|
| 142 |
-
observation = env.reset(task_index=0)
|
| 143 |
-
result = env.step(LookAction(heading_deg=90))
|
| 144 |
-
```
|
| 145 |
-
"""
|
| 146 |
-
|
| 147 |
-
SUPPORTS_CONCURRENT_SESSIONS = True
|
| 148 |
-
|
| 149 |
-
def __init__(
|
| 150 |
-
self,
|
| 151 |
-
index_path: str | None = None,
|
| 152 |
-
cache_dir: str = "",
|
| 153 |
-
episode_mode: str = EpisodeMode.AGENTIC.value,
|
| 154 |
-
max_steps: int = 24,
|
| 155 |
-
max_free_calls: int = 24,
|
| 156 |
-
reward_mode: str = RewardMode.COORDS.value,
|
| 157 |
-
hierarchical_reward: bool = False,
|
| 158 |
-
reward_shape: str = "geoguessr",
|
| 159 |
-
cost_mode: str = "subtract",
|
| 160 |
-
hide_task_identity: bool = False,
|
| 161 |
-
view_size: int = 640,
|
| 162 |
-
allow_fetch: bool = True,
|
| 163 |
-
hires_zoom: bool = True,
|
| 164 |
-
reveal_map: bool = True,
|
| 165 |
-
splits: dict[str, str] | None = None,
|
| 166 |
-
default_split: str = "train",
|
| 167 |
-
):
|
| 168 |
-
if splits:
|
| 169 |
-
self._split_paths = {name: str(path) for name, path in splits.items()}
|
| 170 |
-
elif index_path:
|
| 171 |
-
# Single-index callers keep working: one nameless index becomes the
|
| 172 |
-
# default split, so existing tests, harness runs and the play UI
|
| 173 |
-
# need no change.
|
| 174 |
-
self._split_paths = {default_split: str(index_path)}
|
| 175 |
-
else:
|
| 176 |
-
raise ValueError("Provide either splits= or index_path=.")
|
| 177 |
-
if default_split not in self._split_paths:
|
| 178 |
-
raise ValueError(
|
| 179 |
-
f"default_split {default_split!r} is not one of "
|
| 180 |
-
f"{sorted(self._split_paths)}."
|
| 181 |
-
)
|
| 182 |
-
self._default_split = default_split
|
| 183 |
-
self._cache_dir = cache_dir
|
| 184 |
-
self._allow_fetch = allow_fetch
|
| 185 |
-
self._hires_zoom = hires_zoom
|
| 186 |
-
self._backends: dict[str, PanoramaBackend] = {}
|
| 187 |
-
self._split = default_split
|
| 188 |
-
self._backend = self._backend_for(default_split)
|
| 189 |
-
self._reveal_map = reveal_map
|
| 190 |
-
self._mode = EpisodeMode(episode_mode)
|
| 191 |
-
self._max_steps = max_steps
|
| 192 |
-
# Free tools reveal nothing, so they stay free -- but not unlimited.
|
| 193 |
-
self._max_free_calls = max_free_calls
|
| 194 |
-
self._reward_mode = RewardMode(reward_mode)
|
| 195 |
-
self._hierarchical = hierarchical_reward
|
| 196 |
-
self._reward_shape = reward_shape
|
| 197 |
-
self._cost_mode = cost_mode
|
| 198 |
-
self._hide_task_identity = hide_task_identity
|
| 199 |
-
# Validate now rather than at the end of the first episode.
|
| 200 |
-
compute_reward(1.0, shape=reward_shape, cost_mode=cost_mode)
|
| 201 |
-
self._view_size = (view_size, view_size)
|
| 202 |
-
self._state = GeoGuesserState()
|
| 203 |
-
self._task = None
|
| 204 |
-
self._rng = random.Random()
|
| 205 |
-
|
| 206 |
-
mcp = FastMCP("geoguesser_env")
|
| 207 |
-
self._register_tools(mcp)
|
| 208 |
-
super().__init__(mcp)
|
| 209 |
-
|
| 210 |
-
# -- splits ------------------------------------------------------------
|
| 211 |
-
|
| 212 |
-
def _backend_for(self, split: str) -> PanoramaBackend:
|
| 213 |
-
"""
|
| 214 |
-
Return the backend serving one split, building it on first use.
|
| 215 |
-
|
| 216 |
-
Splits are built lazily so a deployment that only mounts the eval index
|
| 217 |
-
is not forced to carry a training one, and because the parsed index is
|
| 218 |
-
cached process-wide anyway.
|
| 219 |
-
|
| 220 |
-
Args:
|
| 221 |
-
split (`str`):
|
| 222 |
-
Split name.
|
| 223 |
-
|
| 224 |
-
Returns:
|
| 225 |
-
[`PanoramaBackend`]: Backend for that split.
|
| 226 |
-
|
| 227 |
-
Raises:
|
| 228 |
-
UnknownSplitError: When the split is not configured.
|
| 229 |
-
"""
|
| 230 |
-
if split not in self._split_paths:
|
| 231 |
-
raise UnknownSplitError(
|
| 232 |
-
f"Unknown split {split!r}. Available: {sorted(self._split_paths)}."
|
| 233 |
-
)
|
| 234 |
-
backend = self._backends.get(split)
|
| 235 |
-
if backend is None:
|
| 236 |
-
backend = PanoramaBackend(
|
| 237 |
-
self._split_paths[split],
|
| 238 |
-
self._cache_dir,
|
| 239 |
-
allow_fetch=self._allow_fetch,
|
| 240 |
-
hires_zoom=self._hires_zoom,
|
| 241 |
-
)
|
| 242 |
-
self._backends[split] = backend
|
| 243 |
-
return backend
|
| 244 |
-
|
| 245 |
-
@staticmethod
|
| 246 |
-
def _split_type(split: str) -> str:
|
| 247 |
-
"""
|
| 248 |
-
Map a split name onto the type vocabulary the core Task API knows.
|
| 249 |
-
|
| 250 |
-
Core normalises anything outside `{train, validation, test}` to
|
| 251 |
-
`validation`, so `eval` is declared as `test` explicitly rather than
|
| 252 |
-
being silently downgraded.
|
| 253 |
-
"""
|
| 254 |
-
if split == "train":
|
| 255 |
-
return "train"
|
| 256 |
-
if split in {"eval", "test"}:
|
| 257 |
-
return "test"
|
| 258 |
-
return "validation"
|
| 259 |
-
|
| 260 |
-
def _task_spec(self, split: str, task: Task) -> dict[str, Any]:
|
| 261 |
-
"""
|
| 262 |
-
Describe one task for the Task API.
|
| 263 |
-
|
| 264 |
-
Deliberately truth-free: no coordinates and no country. Task specs
|
| 265 |
-
travel to whatever orchestrates training, and the moment a label sits
|
| 266 |
-
in a spec someone can build a prompt from it. The true location is
|
| 267 |
-
revealed in the observation metadata after the guess, which is the one
|
| 268 |
-
place it belongs.
|
| 269 |
-
"""
|
| 270 |
-
return {
|
| 271 |
-
"task_index": task.task_index,
|
| 272 |
-
"task_id": task.task_id,
|
| 273 |
-
"split": split,
|
| 274 |
-
"n_frames": len(task.frames),
|
| 275 |
-
"provider": task.provider,
|
| 276 |
-
"sequence_id": task.sequence_id,
|
| 277 |
-
"offline_ready": bool(task.meta.get("offline_ready", False)),
|
| 278 |
-
}
|
| 279 |
-
|
| 280 |
-
def list_splits(self) -> list[dict[str, Any]]:
|
| 281 |
-
"""
|
| 282 |
-
Task API: describe every configured split.
|
| 283 |
-
|
| 284 |
-
Returns:
|
| 285 |
-
`list[dict]` with keys:
|
| 286 |
-
- `name` (`str`):
|
| 287 |
-
Split name, as accepted by `reset(split=)`.
|
| 288 |
-
- `type` (`str`):
|
| 289 |
-
One of `train`, `test` or `validation`.
|
| 290 |
-
- `num_tasks` (`int`):
|
| 291 |
-
Task count in the split.
|
| 292 |
-
- `default` (`bool`):
|
| 293 |
-
Whether `reset()` uses this split when none is given.
|
| 294 |
-
"""
|
| 295 |
-
return [
|
| 296 |
-
{
|
| 297 |
-
"name": name,
|
| 298 |
-
"type": self._split_type(name),
|
| 299 |
-
"num_tasks": self._backend_for(name).n_tasks,
|
| 300 |
-
"default": name == self._default_split,
|
| 301 |
-
}
|
| 302 |
-
for name in self._split_paths
|
| 303 |
-
]
|
| 304 |
-
|
| 305 |
-
def num_tasks(self, split: str) -> int:
|
| 306 |
-
"""Task API: how many tasks a split holds."""
|
| 307 |
-
return self._backend_for(split).n_tasks
|
| 308 |
-
|
| 309 |
-
def get_task(self, split: str, index: int) -> dict[str, Any]:
|
| 310 |
-
"""Task API: describe one task by split and index."""
|
| 311 |
-
return self._task_spec(split, self._backend_for(split).task(index))
|
| 312 |
-
|
| 313 |
-
def list_tasks(self, split: str) -> list[dict[str, Any]]:
|
| 314 |
-
"""Task API: describe every task in a split."""
|
| 315 |
-
backend = self._backend_for(split)
|
| 316 |
-
return [
|
| 317 |
-
self._task_spec(split, backend.task(index))
|
| 318 |
-
for index in range(backend.n_tasks)
|
| 319 |
-
]
|
| 320 |
-
|
| 321 |
-
def get_task_range(
|
| 322 |
-
self, split: str, start: int | None = None, stop: int | None = None
|
| 323 |
-
) -> list[dict[str, Any]]:
|
| 324 |
-
"""Task API: describe a slice-style range of tasks in a split."""
|
| 325 |
-
backend = self._backend_for(split)
|
| 326 |
-
indices = range(*slice(start, stop).indices(backend.n_tasks))
|
| 327 |
-
return [self._task_spec(split, backend.task(index)) for index in indices]
|
| 328 |
-
|
| 329 |
-
# -- capability-aware tool registration --------------------------------
|
| 330 |
-
|
| 331 |
-
def _navigational(self) -> bool:
|
| 332 |
-
return self._mode is EpisodeMode.AGENTIC and self._backend.supports_move
|
| 333 |
-
|
| 334 |
-
def _can_look(self) -> bool:
|
| 335 |
-
return self._mode is EpisodeMode.AGENTIC and self._backend.supports_look
|
| 336 |
-
|
| 337 |
-
def _can_gather(self) -> bool:
|
| 338 |
-
"""Whether the episode has an investigation phase at all.
|
| 339 |
-
|
| 340 |
-
`single_shot` deliberately has none: one view, one guess, which is the
|
| 341 |
-
shape a VLM GRPO run wants. `nmpz` keeps the map but takes the camera
|
| 342 |
-
away, mirroring the game's own hardest mode.
|
| 343 |
-
"""
|
| 344 |
-
return self._mode is not EpisodeMode.SINGLE_SHOT
|
| 345 |
-
|
| 346 |
-
def _register_tools(self, mcp: FastMCP) -> None:
|
| 347 |
-
"""Register only the tools this configuration can actually serve."""
|
| 348 |
-
if self._can_look():
|
| 349 |
-
|
| 350 |
-
@mcp.tool
|
| 351 |
-
def look(
|
| 352 |
-
heading_deg: float, pitch_deg: float = 0.0, fov_deg: float = 90.0
|
| 353 |
-
) -> str:
|
| 354 |
-
"""Look in a direction. heading_deg is absolute, 0 = true north.
|
| 355 |
-
|
| 356 |
-
Args:
|
| 357 |
-
heading_deg: Compass heading in degrees.
|
| 358 |
-
pitch_deg: Vertical angle; positive looks up.
|
| 359 |
-
fov_deg: Field of view; smaller values zoom in.
|
| 360 |
-
"""
|
| 361 |
-
return self._apply(
|
| 362 |
-
LookAction(
|
| 363 |
-
heading_deg=heading_deg, pitch_deg=pitch_deg, fov_deg=fov_deg
|
| 364 |
-
)
|
| 365 |
-
).feedback
|
| 366 |
-
|
| 367 |
-
@mcp.tool
|
| 368 |
-
def pan(delta_deg: float) -> str:
|
| 369 |
-
"""Turn relative to the current heading; positive turns right.
|
| 370 |
-
|
| 371 |
-
Args:
|
| 372 |
-
delta_deg: Degrees to turn.
|
| 373 |
-
"""
|
| 374 |
-
return self._apply(PanAction(delta_deg=delta_deg)).feedback
|
| 375 |
-
|
| 376 |
-
@mcp.tool
|
| 377 |
-
def zoom(fov_deg: float) -> str:
|
| 378 |
-
"""Change field of view without turning. 30 reads distant signs.
|
| 379 |
-
|
| 380 |
-
Args:
|
| 381 |
-
fov_deg: New field of view in degrees.
|
| 382 |
-
"""
|
| 383 |
-
return self._apply(ZoomAction(fov_deg=fov_deg)).feedback
|
| 384 |
-
|
| 385 |
-
if self._navigational():
|
| 386 |
-
|
| 387 |
-
@mcp.tool
|
| 388 |
-
def move(direction: str, meters: float = 10.0) -> str:
|
| 389 |
-
"""Walk along the road. Reports how far you actually travelled.
|
| 390 |
-
|
| 391 |
-
Args:
|
| 392 |
-
direction: Either "forward" or "backward".
|
| 393 |
-
meters: Requested distance in metres.
|
| 394 |
-
"""
|
| 395 |
-
return self._apply(
|
| 396 |
-
MoveAction(direction=direction, meters=meters)
|
| 397 |
-
).feedback
|
| 398 |
-
|
| 399 |
-
if self._can_gather():
|
| 400 |
-
self._register_map_tools(mcp)
|
| 401 |
-
self._register_guess_tool(mcp)
|
| 402 |
-
|
| 403 |
-
def _register_map_tools(self, mcp: FastMCP) -> None:
|
| 404 |
-
"""Register the map and pin tools, which every gathering mode has."""
|
| 405 |
-
|
| 406 |
-
@mcp.tool
|
| 407 |
-
def place_pin(
|
| 408 |
-
lat: float, lon: float, label: str = "", span_deg: float = 7.0
|
| 409 |
-
) -> str:
|
| 410 |
-
"""Pin a candidate and see where it falls on the map.
|
| 411 |
-
|
| 412 |
-
Tells you what is at that coordinate. Says nothing about whether
|
| 413 |
-
you are right.
|
| 414 |
-
|
| 415 |
-
Args:
|
| 416 |
-
lat: Latitude of the candidate.
|
| 417 |
-
lon: Longitude of the candidate.
|
| 418 |
-
label: Optional note.
|
| 419 |
-
span_deg: Half-width of the map window in degrees. Below about
|
| 420 |
-
4 the map adds roads, urban areas and town names, which is
|
| 421 |
-
how you aim within a city rather than at its centre.
|
| 422 |
-
"""
|
| 423 |
-
return self._apply(
|
| 424 |
-
PinAction(lat=lat, lon=lon, label=label or None, span_deg=span_deg)
|
| 425 |
-
).feedback
|
| 426 |
-
|
| 427 |
-
@mcp.tool
|
| 428 |
-
def view_map(lat: float, lon: float, span_deg: float = 7.0) -> str:
|
| 429 |
-
"""Pan and zoom the map without placing a pin.
|
| 430 |
-
|
| 431 |
-
Args:
|
| 432 |
-
lat: Latitude at the centre of the view.
|
| 433 |
-
lon: Longitude at the centre of the view.
|
| 434 |
-
span_deg: Half-width of the window in degrees.
|
| 435 |
-
"""
|
| 436 |
-
return self._apply(
|
| 437 |
-
ViewMapAction(lat=lat, lon=lon, span_deg=span_deg)
|
| 438 |
-
).feedback
|
| 439 |
-
|
| 440 |
-
@mcp.tool
|
| 441 |
-
def list_pins() -> str:
|
| 442 |
-
"""List the candidates pinned so far. Free."""
|
| 443 |
-
if not self._state.pins:
|
| 444 |
-
return "No pins placed yet."
|
| 445 |
-
return "\n".join(
|
| 446 |
-
f"{i}. {p['lat']:.4f}, {p['lon']:.4f} - {p['description']}"
|
| 447 |
-
for i, p in enumerate(self._state.pins, 1)
|
| 448 |
-
)
|
| 449 |
-
|
| 450 |
-
@mcp.tool
|
| 451 |
-
def clear_pins() -> str:
|
| 452 |
-
"""Remove all pins. Free."""
|
| 453 |
-
self._state.pins = []
|
| 454 |
-
return "Pins cleared."
|
| 455 |
-
|
| 456 |
-
@mcp.tool
|
| 457 |
-
def measure(lat_a: float, lon_a: float, lat_b: float, lon_b: float) -> str:
|
| 458 |
-
"""Distance in km between two coordinates of your own choosing. Free.
|
| 459 |
-
|
| 460 |
-
Args:
|
| 461 |
-
lat_a: Latitude of the first point.
|
| 462 |
-
lon_a: Longitude of the first point.
|
| 463 |
-
lat_b: Latitude of the second point.
|
| 464 |
-
lon_b: Longitude of the second point.
|
| 465 |
-
"""
|
| 466 |
-
km = haversine_km(lat_a, lon_a, lat_b, lon_b)
|
| 467 |
-
return f"{km:.0f} km between those two points."
|
| 468 |
-
|
| 469 |
-
@mcp.tool
|
| 470 |
-
def reverse_geocode(lat: float, lon: float) -> str:
|
| 471 |
-
"""Name the country and nearest city at a coordinate. Free.
|
| 472 |
-
|
| 473 |
-
Args:
|
| 474 |
-
lat: Latitude in degrees.
|
| 475 |
-
lon: Longitude in degrees.
|
| 476 |
-
"""
|
| 477 |
-
place = locate(lat, lon)
|
| 478 |
-
where = place.country or "open water"
|
| 479 |
-
return (
|
| 480 |
-
f"{lat:.4f}, {lon:.4f} is in {where}. Nearest major city: "
|
| 481 |
-
f"{place.nearest_city}, ~{place.city_distance_km:.0f} km "
|
| 482 |
-
f"{place.city_bearing}."
|
| 483 |
-
)
|
| 484 |
-
|
| 485 |
-
def _register_guess_tool(self, mcp: FastMCP) -> None:
|
| 486 |
-
"""Register the terminal action, which every mode has."""
|
| 487 |
-
|
| 488 |
-
@mcp.tool
|
| 489 |
-
def submit_guess(
|
| 490 |
-
lat: float,
|
| 491 |
-
lon: float,
|
| 492 |
-
country: str = "",
|
| 493 |
-
confidence: float = -1.0,
|
| 494 |
-
reasoning: str = "",
|
| 495 |
-
) -> str:
|
| 496 |
-
"""Commit your final answer. Ends the episode.
|
| 497 |
-
|
| 498 |
-
Args:
|
| 499 |
-
lat: Latitude of your guess.
|
| 500 |
-
lon: Longitude of your guess.
|
| 501 |
-
country: Optional ISO-3166 alpha-2 code or country name.
|
| 502 |
-
confidence: Optional self-reported confidence in [0, 1].
|
| 503 |
-
reasoning: Optional rationale, recorded but not scored.
|
| 504 |
-
"""
|
| 505 |
-
observation = self._apply(
|
| 506 |
-
GuessAction(
|
| 507 |
-
lat=lat,
|
| 508 |
-
lon=lon,
|
| 509 |
-
country=country or None,
|
| 510 |
-
confidence=None if confidence < 0 else confidence,
|
| 511 |
-
reasoning=reasoning or None,
|
| 512 |
-
)
|
| 513 |
-
)
|
| 514 |
-
return observation.feedback
|
| 515 |
-
|
| 516 |
-
# -- lifecycle ---------------------------------------------------------
|
| 517 |
-
|
| 518 |
-
def reset(
|
| 519 |
-
self,
|
| 520 |
-
seed: int | None = None,
|
| 521 |
-
episode_id: str | None = None,
|
| 522 |
-
split: str | None = None,
|
| 523 |
-
index: int | None = None,
|
| 524 |
-
task_index: int | None = None,
|
| 525 |
-
**kwargs: Any,
|
| 526 |
-
) -> GeoGuesserObservation:
|
| 527 |
-
"""
|
| 528 |
-
Start an episode.
|
| 529 |
-
|
| 530 |
-
Selection is explicit, because training and demoing want opposite
|
| 531 |
-
things. `index` picks one exact task and is byte-identical on repeat,
|
| 532 |
-
which is what a GRPO group needs. `seed` picks `tasks[seed % n_tasks]`.
|
| 533 |
-
Neither means a random task; omitting both does, and the chosen split
|
| 534 |
-
and index are always reported in the observation metadata so a random
|
| 535 |
-
episode stays replayable.
|
| 536 |
-
|
| 537 |
-
Args:
|
| 538 |
-
seed (`int`, *optional*):
|
| 539 |
-
Deterministic selector, `tasks[seed % n_tasks]`.
|
| 540 |
-
episode_id (`str`, *optional*):
|
| 541 |
-
Caller-supplied episode identifier.
|
| 542 |
-
split (`str`, *optional*):
|
| 543 |
-
Which split to draw from. Defaults to the environment's default
|
| 544 |
-
split.
|
| 545 |
-
index (`int`, *optional*):
|
| 546 |
-
Exact task to play, within `split`. Takes precedence over
|
| 547 |
-
`seed`.
|
| 548 |
-
task_index (`int`, *optional*):
|
| 549 |
-
Deprecated alias for `index`, kept so existing callers and
|
| 550 |
-
saved trajectories keep working.
|
| 551 |
-
|
| 552 |
-
Returns:
|
| 553 |
-
[`GeoGuesserObservation`]: The opening view and prompt.
|
| 554 |
-
|
| 555 |
-
Raises:
|
| 556 |
-
UnknownSplitError: When `split` is not configured.
|
| 557 |
-
"""
|
| 558 |
-
self._split = split or self._default_split
|
| 559 |
-
self._backend = self._backend_for(self._split)
|
| 560 |
-
|
| 561 |
-
if index is None:
|
| 562 |
-
index = task_index
|
| 563 |
-
n = self._backend.n_tasks
|
| 564 |
-
if index is not None:
|
| 565 |
-
chosen = int(index) % n
|
| 566 |
-
elif seed is not None:
|
| 567 |
-
chosen = int(seed) % n
|
| 568 |
-
else:
|
| 569 |
-
chosen = self._rng.randrange(n)
|
| 570 |
-
|
| 571 |
-
self._task = self._backend.task(chosen)
|
| 572 |
-
start = self._task.frames[self._task.start_frame]
|
| 573 |
-
self._state = GeoGuesserState(
|
| 574 |
-
episode_id=episode_id or str(uuid.uuid4()),
|
| 575 |
-
task_index=chosen,
|
| 576 |
-
task_id=self._task.task_id,
|
| 577 |
-
frame_index=self._task.start_frame,
|
| 578 |
-
heading_deg=0.0,
|
| 579 |
-
pitch_deg=0.0,
|
| 580 |
-
fov_deg=90.0,
|
| 581 |
-
)
|
| 582 |
-
|
| 583 |
-
observation = self._render_view_observation()
|
| 584 |
-
observation.prompt = PROMPT.format(
|
| 585 |
-
tools=", ".join(self._tool_names()), steps=self._max_steps
|
| 586 |
-
)
|
| 587 |
-
observation.feedback = "Episode started."
|
| 588 |
-
observation.captured_at = start.captured_at
|
| 589 |
-
observation.metadata = self._metadata()
|
| 590 |
-
return observation
|
| 591 |
-
|
| 592 |
-
@property
|
| 593 |
-
def state(self) -> GeoGuesserState:
|
| 594 |
-
"""Current internal state."""
|
| 595 |
-
return self._state
|
| 596 |
-
|
| 597 |
-
def _step_impl(self, action: Action, **kwargs: Any) -> Observation:
|
| 598 |
-
"""Handle structured, non-MCP actions.
|
| 599 |
-
|
| 600 |
-
Accepts either the flat wire action the HTTP layer delivers or a typed
|
| 601 |
-
action constructed in process, so tests and the harness can bypass
|
| 602 |
-
serialisation without a second code path.
|
| 603 |
-
"""
|
| 604 |
-
if isinstance(action, GeoGuesserAction):
|
| 605 |
-
return self._apply(from_wire(action))
|
| 606 |
-
if isinstance(action, TypedAction):
|
| 607 |
-
return self._apply(action)
|
| 608 |
-
raise TypeError(f"Unsupported action type: {type(action).__name__}")
|
| 609 |
-
|
| 610 |
-
# -- the actual mechanics ---------------------------------------------
|
| 611 |
-
|
| 612 |
-
def _tool_names(self) -> list[str]:
|
| 613 |
-
"""Names of the tools this configuration actually registered."""
|
| 614 |
-
names: list[str] = []
|
| 615 |
-
if self._can_look():
|
| 616 |
-
names += ["look", "pan", "zoom"]
|
| 617 |
-
if self._navigational():
|
| 618 |
-
names.append("move")
|
| 619 |
-
if self._can_gather():
|
| 620 |
-
names += [
|
| 621 |
-
"place_pin",
|
| 622 |
-
"view_map",
|
| 623 |
-
"list_pins",
|
| 624 |
-
"clear_pins",
|
| 625 |
-
"measure",
|
| 626 |
-
"reverse_geocode",
|
| 627 |
-
]
|
| 628 |
-
names.append("submit_guess")
|
| 629 |
-
return names
|
| 630 |
-
|
| 631 |
-
def _steps_used(self) -> int:
|
| 632 |
-
s = self._state
|
| 633 |
-
return s.n_looks + s.n_maps + s.n_pins + s.n_moves
|
| 634 |
-
|
| 635 |
-
def _steps_remaining(self) -> int:
|
| 636 |
-
return max(0, self._max_steps - self._steps_used())
|
| 637 |
-
|
| 638 |
-
def _cost(self) -> float:
|
| 639 |
-
s = self._state
|
| 640 |
-
return action_cost(
|
| 641 |
-
n_looks=s.n_looks, n_maps=s.n_maps, n_pins=s.n_pins, n_moves=s.n_moves
|
| 642 |
-
)
|
| 643 |
-
|
| 644 |
-
def _metadata(self) -> dict[str, Any]:
|
| 645 |
-
# Identity fields are a reward-hacking channel, not just clutter. The
|
| 646 |
-
# Mapillary contributor determines the country outright for 74% of
|
| 647 |
-
# training tasks ("amsterdam" only maps the Netherlands), and
|
| 648 |
-
# task_index/task_id/sequence_id are a few thousand memorisable keys
|
| 649 |
-
# straight to a coordinate -- either lets a policy score without ever
|
| 650 |
-
# reading the image. Kept by default so the play UI can show the licence
|
| 651 |
-
# credit and recordings keep their provenance; RL training must set
|
| 652 |
-
# hide_task_identity=True. The terminal observation carries them
|
| 653 |
-
# regardless, since by then the truth is already revealed.
|
| 654 |
-
if self._hide_task_identity:
|
| 655 |
-
return {
|
| 656 |
-
"split": self._split,
|
| 657 |
-
"street_detail": (
|
| 658 |
-
"unavailable"
|
| 659 |
-
if street_fetch_failed()
|
| 660 |
-
else ("on" if street_detail_enabled() else "off")
|
| 661 |
-
),
|
| 662 |
-
"backend": "mapillary",
|
| 663 |
-
"episode_mode": self._mode.value,
|
| 664 |
-
"frame_index": self._state.frame_index,
|
| 665 |
-
}
|
| 666 |
-
return {
|
| 667 |
-
"split": self._split,
|
| 668 |
-
# A map quietly missing its streets looks like a styling choice, so
|
| 669 |
-
# say so. Overpass 504s from datacenter egress, which is how a Space
|
| 670 |
-
# ends up with poorer maps than a laptop for the same task.
|
| 671 |
-
"street_detail": (
|
| 672 |
-
"unavailable"
|
| 673 |
-
if street_fetch_failed()
|
| 674 |
-
else ("on" if street_detail_enabled() else "off")
|
| 675 |
-
),
|
| 676 |
-
"task_index": self._state.task_index,
|
| 677 |
-
"task_id": self._state.task_id,
|
| 678 |
-
"backend": "mapillary",
|
| 679 |
-
# Opaque upstream identifier, not ground truth: it makes a recorded
|
| 680 |
-
# episode traceable back to its source sequence.
|
| 681 |
-
"sequence_id": self._task.sequence_id if self._task else None,
|
| 682 |
-
"episode_mode": self._mode.value,
|
| 683 |
-
"frame_index": self._state.frame_index,
|
| 684 |
-
"captured_at": self._task.frames[self._state.frame_index].captured_at,
|
| 685 |
-
"attribution": self._task.attribution,
|
| 686 |
-
}
|
| 687 |
-
|
| 688 |
-
def _base_observation(self) -> GeoGuesserObservation:
|
| 689 |
-
s = self._state
|
| 690 |
-
return GeoGuesserObservation(
|
| 691 |
-
heading_deg=s.heading_deg % 360,
|
| 692 |
-
pitch_deg=s.pitch_deg,
|
| 693 |
-
fov_deg=s.fov_deg,
|
| 694 |
-
total_moved_meters=s.total_moved_meters,
|
| 695 |
-
can_move_forward=self._navigational()
|
| 696 |
-
and self._backend.can_move(self._task, s.frame_index, "forward"),
|
| 697 |
-
can_move_backward=self._navigational()
|
| 698 |
-
and self._backend.can_move(self._task, s.frame_index, "backward"),
|
| 699 |
-
available_tools=self._tool_names(),
|
| 700 |
-
steps_remaining=self._steps_remaining(),
|
| 701 |
-
action_cost=self._cost(),
|
| 702 |
-
pins=[Pin(**p) for p in s.pins],
|
| 703 |
-
captured_at=self._task.frames[s.frame_index].captured_at,
|
| 704 |
-
metadata=self._metadata(),
|
| 705 |
-
)
|
| 706 |
-
|
| 707 |
-
def _render_view_observation(self) -> GeoGuesserObservation:
|
| 708 |
-
s = self._state
|
| 709 |
-
view = self._backend.render_view(
|
| 710 |
-
self._task, s.frame_index, s.heading_deg, s.pitch_deg, s.fov_deg
|
| 711 |
-
)
|
| 712 |
-
if view.size != self._view_size:
|
| 713 |
-
view = view.resize(self._view_size)
|
| 714 |
-
observation = self._base_observation()
|
| 715 |
-
observation.image_base64 = to_base64(view, "JPEG")
|
| 716 |
-
observation.image_kind = "view"
|
| 717 |
-
return observation
|
| 718 |
-
|
| 719 |
-
def _render_map_observation(
|
| 720 |
-
self,
|
| 721 |
-
pins: list[tuple[float, float]],
|
| 722 |
-
focus,
|
| 723 |
-
span: float,
|
| 724 |
-
truth: tuple[float, float] | None = None,
|
| 725 |
-
) -> GeoGuesserObservation:
|
| 726 |
-
image = render_map(pins, focus=focus, span_deg=span, truth=truth)
|
| 727 |
-
observation = self._base_observation()
|
| 728 |
-
observation.image_base64 = to_base64(image, "PNG")
|
| 729 |
-
observation.image_kind = "map"
|
| 730 |
-
return observation
|
| 731 |
-
|
| 732 |
-
def _apply(self, action: Action) -> GeoGuesserObservation:
|
| 733 |
-
"""Execute one action and produce the resulting observation."""
|
| 734 |
-
if self._task is None:
|
| 735 |
-
raise RuntimeError("reset() must be called before step().")
|
| 736 |
-
|
| 737 |
-
s = self._state
|
| 738 |
-
s.step_count += 1
|
| 739 |
-
|
| 740 |
-
if s.submitted:
|
| 741 |
-
observation = self._base_observation()
|
| 742 |
-
observation.done = True
|
| 743 |
-
observation.feedback = "The episode is over; the guess was already made."
|
| 744 |
-
return observation
|
| 745 |
-
|
| 746 |
-
if isinstance(action, GuessAction):
|
| 747 |
-
return self._finish(action)
|
| 748 |
-
|
| 749 |
-
if self._steps_remaining() <= 0:
|
| 750 |
-
observation = self._base_observation()
|
| 751 |
-
observation.feedback = (
|
| 752 |
-
"Out of actions. Call submit_guess with your best estimate."
|
| 753 |
-
)
|
| 754 |
-
return observation
|
| 755 |
-
|
| 756 |
-
if isinstance(action, LookAction):
|
| 757 |
-
s.heading_deg = action.heading_deg
|
| 758 |
-
s.pitch_deg = action.pitch_deg
|
| 759 |
-
s.fov_deg = action.fov_deg
|
| 760 |
-
s.n_looks += 1
|
| 761 |
-
observation = self._render_view_observation()
|
| 762 |
-
observation.feedback = (
|
| 763 |
-
f"Facing {s.heading_deg % 360:.0f} deg, {s.fov_deg:.0f} deg field of "
|
| 764 |
-
f"view. {self._steps_remaining()} actions left."
|
| 765 |
-
)
|
| 766 |
-
return observation
|
| 767 |
-
|
| 768 |
-
if isinstance(action, PanAction):
|
| 769 |
-
s.heading_deg = (s.heading_deg + action.delta_deg) % 360
|
| 770 |
-
s.n_looks += 1
|
| 771 |
-
observation = self._render_view_observation()
|
| 772 |
-
observation.feedback = (
|
| 773 |
-
f"Turned to {s.heading_deg:.0f} deg. "
|
| 774 |
-
f"{self._steps_remaining()} actions left."
|
| 775 |
-
)
|
| 776 |
-
return observation
|
| 777 |
-
|
| 778 |
-
if isinstance(action, ZoomAction):
|
| 779 |
-
s.fov_deg = action.fov_deg
|
| 780 |
-
s.n_looks += 1
|
| 781 |
-
observation = self._render_view_observation()
|
| 782 |
-
observation.feedback = (
|
| 783 |
-
f"Field of view now {s.fov_deg:.0f} deg. "
|
| 784 |
-
f"{self._steps_remaining()} actions left."
|
| 785 |
-
)
|
| 786 |
-
return observation
|
| 787 |
-
|
| 788 |
-
if isinstance(action, MoveAction):
|
| 789 |
-
new_index, travelled = self._backend.step_along(
|
| 790 |
-
self._task, s.frame_index, action.direction, action.meters
|
| 791 |
-
)
|
| 792 |
-
s.n_moves += 1
|
| 793 |
-
if new_index == s.frame_index:
|
| 794 |
-
observation = self._render_view_observation()
|
| 795 |
-
observation.feedback = (
|
| 796 |
-
f"Cannot go {action.direction} - the captured road ends here. "
|
| 797 |
-
f"{self._steps_remaining()} actions left."
|
| 798 |
-
)
|
| 799 |
-
return observation
|
| 800 |
-
s.frame_index = new_index
|
| 801 |
-
s.total_moved_meters += travelled
|
| 802 |
-
observation = self._render_view_observation()
|
| 803 |
-
observation.moved_meters = travelled
|
| 804 |
-
observation.feedback = (
|
| 805 |
-
f"Moved {travelled:.0f} m {action.direction} "
|
| 806 |
-
f"({s.total_moved_meters:.0f} m total). "
|
| 807 |
-
f"{self._steps_remaining()} actions left."
|
| 808 |
-
)
|
| 809 |
-
return observation
|
| 810 |
-
|
| 811 |
-
if isinstance(action, PinAction):
|
| 812 |
-
previous = (s.pins[-1]["lat"], s.pins[-1]["lon"]) if s.pins else None
|
| 813 |
-
description = describe_pin(
|
| 814 |
-
len(s.pins) + 1, action.lat, action.lon, previous=previous
|
| 815 |
-
)
|
| 816 |
-
s.pins.append(
|
| 817 |
-
{
|
| 818 |
-
"index": len(s.pins) + 1,
|
| 819 |
-
"lat": action.lat,
|
| 820 |
-
"lon": action.lon,
|
| 821 |
-
"label": action.label,
|
| 822 |
-
"description": description,
|
| 823 |
-
}
|
| 824 |
-
)
|
| 825 |
-
s.n_pins += 1
|
| 826 |
-
pins = [(p["lat"], p["lon"]) for p in s.pins]
|
| 827 |
-
observation = self._render_map_observation(
|
| 828 |
-
pins, (action.lat, action.lon), action.span_deg
|
| 829 |
-
)
|
| 830 |
-
observation.feedback = (
|
| 831 |
-
f"{description} {self._steps_remaining()} actions left."
|
| 832 |
-
)
|
| 833 |
-
return observation
|
| 834 |
-
|
| 835 |
-
if isinstance(action, ViewMapAction):
|
| 836 |
-
s.n_maps += 1
|
| 837 |
-
pins = [(p["lat"], p["lon"]) for p in s.pins]
|
| 838 |
-
observation = self._render_map_observation(
|
| 839 |
-
pins, (action.lat, action.lon), action.span_deg
|
| 840 |
-
)
|
| 841 |
-
place = locate(action.lat, action.lon)
|
| 842 |
-
observation.feedback = (
|
| 843 |
-
f"Map centred on {action.lat:.3f}, {action.lon:.3f} "
|
| 844 |
-
f"({place.country or 'open water'}), "
|
| 845 |
-
f"{action.span_deg * 2:.0f} deg across. "
|
| 846 |
-
f"{self._steps_remaining()} actions left."
|
| 847 |
-
)
|
| 848 |
-
return observation
|
| 849 |
-
|
| 850 |
-
if isinstance(action, MeasureAction):
|
| 851 |
-
# Free, because it is arithmetic on two coordinates the agent
|
| 852 |
-
# supplied: it reveals nothing about where the agent is. But free
|
| 853 |
-
# used to mean unbounded -- it incremented no counter, so it never
|
| 854 |
-
# advanced the step budget and a policy could issue it forever. Past
|
| 855 |
-
# the cap it starts costing a map action, so the episode terminates.
|
| 856 |
-
if s.n_free >= self._max_free_calls:
|
| 857 |
-
s.n_maps += 1
|
| 858 |
-
else:
|
| 859 |
-
s.n_free += 1
|
| 860 |
-
km = haversine_km(action.lat_a, action.lon_a, action.lat_b, action.lon_b)
|
| 861 |
-
observation = self._base_observation()
|
| 862 |
-
observation.feedback = f"{km:.0f} km between those two points."
|
| 863 |
-
if s.n_free >= self._max_free_calls:
|
| 864 |
-
observation.feedback += " Free-tool budget spent; this now costs."
|
| 865 |
-
return observation
|
| 866 |
-
|
| 867 |
-
raise TypeError(f"Unsupported action type: {type(action).__name__}")
|
| 868 |
-
|
| 869 |
-
def _finish(self, action: GuessAction) -> GeoGuesserObservation:
|
| 870 |
-
"""Score the final guess and end the episode."""
|
| 871 |
-
s = self._state
|
| 872 |
-
s.submitted = True
|
| 873 |
-
true_lat, true_lon = self._task.truth
|
| 874 |
-
|
| 875 |
-
if action.lat is not None and action.lon is not None:
|
| 876 |
-
lat, lon, parsed_ok, note = action.lat, action.lon, True, ""
|
| 877 |
-
else:
|
| 878 |
-
parsed = parse_guess(action.response or "")
|
| 879 |
-
lat, lon, parsed_ok, note = parsed.lat, parsed.lon, parsed.ok, parsed.note
|
| 880 |
-
|
| 881 |
-
cost = self._cost()
|
| 882 |
-
observation = self._base_observation()
|
| 883 |
-
observation.done = True
|
| 884 |
-
observation.parsed_ok = parsed_ok
|
| 885 |
-
observation.true_lat = true_lat
|
| 886 |
-
observation.true_lon = true_lon
|
| 887 |
-
observation.action_cost = cost
|
| 888 |
-
|
| 889 |
-
if not parsed_ok:
|
| 890 |
-
observation.reward = 0.0
|
| 891 |
-
observation.score = 0.0
|
| 892 |
-
observation.feedback = f"No usable guess. {note} Scored 0."
|
| 893 |
-
observation.metadata = {
|
| 894 |
-
**self._metadata(),
|
| 895 |
-
"parse_failure": True,
|
| 896 |
-
"country": self._task.country,
|
| 897 |
-
"task_index": self._state.task_index,
|
| 898 |
-
"task_id": self._state.task_id,
|
| 899 |
-
"sequence_id": self._task.sequence_id,
|
| 900 |
-
"attribution": self._task.attribution,
|
| 901 |
-
}
|
| 902 |
-
return observation
|
| 903 |
-
|
| 904 |
-
distance = haversine_km(lat, lon, true_lat, true_lon)
|
| 905 |
-
|
| 906 |
-
# A guess previously returned no image at all, which left the outcome
|
| 907 |
-
# invisible in a trace and gave a policy nothing to learn the shape of
|
| 908 |
-
# its error from. Truth is only ever drawn here, after scoring.
|
| 909 |
-
if not self._reveal_map:
|
| 910 |
-
observation.image_kind = "none"
|
| 911 |
-
separation = max(abs(lat - true_lat), abs(lon - true_lon))
|
| 912 |
-
if self._reveal_map and separation > 25.0:
|
| 913 |
-
# Framing both points would squash a hemisphere into the panel and
|
| 914 |
-
# tell you nothing. The useful second view is where it actually was.
|
| 915 |
-
reveal_focus = (true_lat, true_lon)
|
| 916 |
-
reveal_span = 12.0
|
| 917 |
-
elif self._reveal_map:
|
| 918 |
-
reveal_focus = ((lat + true_lat) / 2, (lon + true_lon) / 2)
|
| 919 |
-
reveal_span = max(0.05, separation * 0.75 + 0.4)
|
| 920 |
-
if self._reveal_map:
|
| 921 |
-
reveal = self._render_map_observation(
|
| 922 |
-
[(lat, lon)], reveal_focus, reveal_span, truth=(true_lat, true_lon)
|
| 923 |
-
)
|
| 924 |
-
observation.image_base64 = reveal.image_base64
|
| 925 |
-
observation.image_kind = "map"
|
| 926 |
-
|
| 927 |
-
truth_place = locate(true_lat, true_lon)
|
| 928 |
-
guess_place = locate(lat, lon)
|
| 929 |
-
country_hit = bool(
|
| 930 |
-
truth_place.country and truth_place.country == guess_place.country
|
| 931 |
-
)
|
| 932 |
-
region_hit = bool(
|
| 933 |
-
truth_place.subregion and truth_place.subregion == guess_place.subregion
|
| 934 |
-
)
|
| 935 |
-
|
| 936 |
-
if self._reward_mode is RewardMode.COUNTRY_ONLY:
|
| 937 |
-
reward = max(0.0, float(country_hit) - cost)
|
| 938 |
-
score = float(country_hit)
|
| 939 |
-
else:
|
| 940 |
-
score = compute_reward(
|
| 941 |
-
distance,
|
| 942 |
-
cost=0.0,
|
| 943 |
-
country_hit=country_hit,
|
| 944 |
-
region_hit=region_hit,
|
| 945 |
-
hierarchical=self._hierarchical,
|
| 946 |
-
shape=self._reward_shape,
|
| 947 |
-
)
|
| 948 |
-
reward = compute_reward(
|
| 949 |
-
distance,
|
| 950 |
-
cost=cost,
|
| 951 |
-
country_hit=country_hit,
|
| 952 |
-
region_hit=region_hit,
|
| 953 |
-
hierarchical=self._hierarchical,
|
| 954 |
-
shape=self._reward_shape,
|
| 955 |
-
cost_mode=self._cost_mode,
|
| 956 |
-
)
|
| 957 |
-
|
| 958 |
-
observation.distance_km = distance
|
| 959 |
-
observation.score = score
|
| 960 |
-
observation.reward = reward
|
| 961 |
-
observation.feedback = (
|
| 962 |
-
f"{verdict(distance)} - {distance:.0f} km away. True location "
|
| 963 |
-
f"{true_lat:.4f}, {true_lon:.4f} "
|
| 964 |
-
f"({truth_place.country or 'open water'}). "
|
| 965 |
-
+ (
|
| 966 |
-
f"Score {score:.3f} scaled by {1 - min(cost, 0.5):.2f} "
|
| 967 |
-
f"action cost = {reward:.3f}."
|
| 968 |
-
if self._cost_mode == "multiply"
|
| 969 |
-
else f"Score {score:.3f} minus {cost:.2f} action cost = {reward:.3f}."
|
| 970 |
-
)
|
| 971 |
-
)
|
| 972 |
-
observation.metadata = {
|
| 973 |
-
**self._metadata(),
|
| 974 |
-
# Revealed here and only here, next to true_lat/true_lon: the
|
| 975 |
-
# country is the label a per-region breakdown needs, and a task spec
|
| 976 |
-
# deliberately does not carry it.
|
| 977 |
-
"country": self._task.country,
|
| 978 |
-
# Restored here even under hide_task_identity: the episode is over,
|
| 979 |
-
# so provenance can no longer be used to shortcut it.
|
| 980 |
-
"task_index": self._state.task_index,
|
| 981 |
-
"task_id": self._state.task_id,
|
| 982 |
-
"sequence_id": self._task.sequence_id,
|
| 983 |
-
"attribution": self._task.attribution,
|
| 984 |
-
"guess": [lat, lon],
|
| 985 |
-
"guess_country": guess_place.country,
|
| 986 |
-
"verdict": verdict(distance),
|
| 987 |
-
"distance_km": distance,
|
| 988 |
-
"country_hit": country_hit,
|
| 989 |
-
"region_hit": region_hit,
|
| 990 |
-
"confidence": action.confidence,
|
| 991 |
-
"reasoning": action.reasoning,
|
| 992 |
-
"n_looks": s.n_looks,
|
| 993 |
-
"n_maps": s.n_maps,
|
| 994 |
-
"n_pins": s.n_pins,
|
| 995 |
-
"n_moves": s.n_moves,
|
| 996 |
-
"total_moved_meters": s.total_moved_meters,
|
| 997 |
-
}
|
| 998 |
-
return observation
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/gradio_ui.py
DELETED
|
@@ -1,1044 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""The human-play page: a GeoGuessr-style game in the browser.
|
| 4 |
-
|
| 5 |
-
Two viewers, both driven by the environment's own data:
|
| 6 |
-
|
| 7 |
-
- Pannellum shows the equirectangular panorama, so a person drags to look
|
| 8 |
-
around exactly where the agent calls `look()`.
|
| 9 |
-
- MapLibre shows the guess map over OpenFreeMap tiles. No API key, no request
|
| 10 |
-
limits, commercial use permitted, and self-hostable if the public instance
|
| 11 |
-
ever goes away.
|
| 12 |
-
|
| 13 |
-
The page is served as its own document at `/geoguesser/play` and embedded in
|
| 14 |
-
the Gradio tab through an iframe, because `gr.HTML` inserts markup without
|
| 15 |
-
executing `<script>` tags — styles apply, but neither viewer initialises, which
|
| 16 |
-
looks like a blank panel and reports no error anywhere.
|
| 17 |
-
|
| 18 |
-
A human sees live tiles; the agent's map stays the offline Natural Earth
|
| 19 |
-
render. They agree on geometry, which is what matters, and the agent keeps a
|
| 20 |
-
determinism the browser does not need.
|
| 21 |
-
"""
|
| 22 |
-
|
| 23 |
-
from __future__ import annotations
|
| 24 |
-
|
| 25 |
-
import json
|
| 26 |
-
import random
|
| 27 |
-
import urllib.parse
|
| 28 |
-
from typing import Any, Dict, List, Optional
|
| 29 |
-
|
| 30 |
-
import gradio as gr
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
MAPLIBRE_JS = "https://cdnjs.cloudflare.com/ajax/libs/maplibre-gl/5.24.0/maplibre-gl.js"
|
| 34 |
-
MAPLIBRE_CSS = (
|
| 35 |
-
"https://cdnjs.cloudflare.com/ajax/libs/maplibre-gl/5.24.0/maplibre-gl.css"
|
| 36 |
-
)
|
| 37 |
-
PANNELLUM_JS = "https://cdnjs.cloudflare.com/ajax/libs/pannellum/2.5.6/pannellum.js"
|
| 38 |
-
PANNELLUM_CSS = "https://cdnjs.cloudflare.com/ajax/libs/pannellum/2.5.6/pannellum.css"
|
| 39 |
-
OPENFREEMAP_STYLE = "https://tiles.openfreemap.org/styles/positron"
|
| 40 |
-
|
| 41 |
-
# The real game scores a round out of 5000 on this curve. Showing points
|
| 42 |
-
# rather than the RL reward makes a score comparable to GeoGuessr intuition;
|
| 43 |
-
# the reward is shown beside it so the two are never confused. There is no
|
| 44 |
-
# multi-round game here, because an episode is exactly one guess.
|
| 45 |
-
MAX_POINTS_PER_ROUND = 5000
|
| 46 |
-
|
| 47 |
-
_TEMPLATE = r"""<!doctype html>
|
| 48 |
-
<html lang="en">
|
| 49 |
-
<head>
|
| 50 |
-
<meta charset="utf-8">
|
| 51 |
-
<meta name="viewport" content="width=device-width,initial-scale=1">
|
| 52 |
-
<title>geoguesser_env - play</title>
|
| 53 |
-
<link rel="stylesheet" href="__MAPLIBRE_CSS__">
|
| 54 |
-
<link rel="stylesheet" href="__PANNELLUM_CSS__">
|
| 55 |
-
<style>
|
| 56 |
-
:root {
|
| 57 |
-
--panel: rgba(18, 23, 28, .92);
|
| 58 |
-
--edge: #2c353d;
|
| 59 |
-
--ink: #e8ecef;
|
| 60 |
-
--ink-soft: #9aa5ad;
|
| 61 |
-
--ink-faint: #6e7a83;
|
| 62 |
-
--pin: #c4332a;
|
| 63 |
-
--good: #5aa06e;
|
| 64 |
-
--wire: #7fb3cc;
|
| 65 |
-
}
|
| 66 |
-
* { box-sizing: border-box; }
|
| 67 |
-
html, body {
|
| 68 |
-
margin: 0; height: 100%; overflow: hidden; background: #10151a;
|
| 69 |
-
color: var(--ink);
|
| 70 |
-
font-family: ui-monospace, "SF Mono", "IBM Plex Mono", Menlo, monospace;
|
| 71 |
-
}
|
| 72 |
-
#stage { position: absolute; inset: 0; }
|
| 73 |
-
#pano { position: absolute; inset: 0; }
|
| 74 |
-
.pnlm-zoom-controls, .pnlm-orientation-button, .pnlm-panorama-info,
|
| 75 |
-
.pnlm-compass { display: none !important; }
|
| 76 |
-
.pnlm-load-box { background: #10151a !important; }
|
| 77 |
-
|
| 78 |
-
.hud {
|
| 79 |
-
position: absolute; z-index: 5; background: var(--panel);
|
| 80 |
-
border: 1px solid var(--edge); border-radius: 4px;
|
| 81 |
-
font-size: 12px; padding: 7px 11px; backdrop-filter: blur(8px);
|
| 82 |
-
line-height: 1.5;
|
| 83 |
-
}
|
| 84 |
-
.hud b { color: #fff; font-weight: 500; }
|
| 85 |
-
.hud span { color: var(--ink-soft); }
|
| 86 |
-
#top { top: 12px; left: 12px; }
|
| 87 |
-
#top .ep { color: var(--wire); }
|
| 88 |
-
#compass { top: 12px; left: 50%; transform: translateX(-50%); letter-spacing: .1em; }
|
| 89 |
-
#score { top: 12px; right: 12px; text-align: right; }
|
| 90 |
-
#score .pts { font-size: 15px; color: #fff; }
|
| 91 |
-
#credit {
|
| 92 |
-
top: 74px; right: 12px; font-size: 10.5px; color: var(--ink-soft);
|
| 93 |
-
max-width: 34vw; text-align: right; z-index: 14;
|
| 94 |
-
}
|
| 95 |
-
#credit a { color: #8fb8cc; text-decoration: none; }
|
| 96 |
-
#actions {
|
| 97 |
-
bottom: 12px; left: 50%; transform: translateX(-50%); display: flex;
|
| 98 |
-
gap: 6px; align-items: center; transition: opacity .3s ease;
|
| 99 |
-
}
|
| 100 |
-
#actions {
|
| 101 |
-
gap: 0; padding: 0; display: flex; align-items: stretch; overflow: hidden;
|
| 102 |
-
bottom: 12px; left: 12px; transform: none;
|
| 103 |
-
}
|
| 104 |
-
/* One group per kind of environment action, each labelled, each showing what
|
| 105 |
-
it costs through its tooltip rather than shouting a number. A group whose
|
| 106 |
-
tool is not registered is removed rather than greyed: a control you cannot
|
| 107 |
-
use is noise. */
|
| 108 |
-
.pad {
|
| 109 |
-
display: flex; flex-direction: column; gap: 4px; padding: 8px 14px;
|
| 110 |
-
justify-content: center;
|
| 111 |
-
}
|
| 112 |
-
.pad + .pad { border-left: 1px solid var(--edge); }
|
| 113 |
-
.pad.gone { display: none; }
|
| 114 |
-
.padlabel {
|
| 115 |
-
font-size: 9.5px; letter-spacing: .14em; text-transform: uppercase;
|
| 116 |
-
color: var(--ink-faint);
|
| 117 |
-
}
|
| 118 |
-
.padlabel b { color: var(--ink); font-weight: 500; letter-spacing: 0; }
|
| 119 |
-
.btns { display: flex; align-items: center; gap: 5px; }
|
| 120 |
-
#actions button {
|
| 121 |
-
font-family: inherit; cursor: pointer; color: var(--ink);
|
| 122 |
-
background: #232c33; border: 1px solid var(--edge); border-radius: 4px;
|
| 123 |
-
display: flex; flex-direction: column; align-items: center; gap: 1px;
|
| 124 |
-
min-width: 46px; padding: 5px 7px; line-height: 1;
|
| 125 |
-
}
|
| 126 |
-
#actions button .glyph { font-size: 12px; }
|
| 127 |
-
#actions button .tag {
|
| 128 |
-
font-size: 8.5px; letter-spacing: .06em; color: var(--ink-soft);
|
| 129 |
-
}
|
| 130 |
-
#actions button:hover {
|
| 131 |
-
border-color: #6d8493; background: #2b353d; color: #fff;
|
| 132 |
-
}
|
| 133 |
-
#actions button:hover .tag { color: var(--ink); }
|
| 134 |
-
#actions button:active { background: #1c242a; }
|
| 135 |
-
#actions button:focus-visible { outline: 2px solid var(--wire); outline-offset: 1px; }
|
| 136 |
-
#actions button.gone { display: none; }
|
| 137 |
-
#actions.working { opacity: .55; }
|
| 138 |
-
#actions.working button { cursor: progress; }
|
| 139 |
-
#actions.working #padBudget b { color: var(--wire); }
|
| 140 |
-
#actions button.preset {
|
| 141 |
-
font-size: 10px; letter-spacing: .04em; min-width: 44px; padding: 7px 8px;
|
| 142 |
-
}
|
| 143 |
-
#actions button.preset.on { border-color: var(--wire); color: #fff; }
|
| 144 |
-
#fovRange { width: 104px; accent-color: var(--wire); margin-left: 4px; }
|
| 145 |
-
.budget { gap: 4px; font-size: 10px; color: var(--ink-faint); white-space: nowrap; }
|
| 146 |
-
.budget .sep { color: var(--edge); margin: 0 3px; }
|
| 147 |
-
.budget b {
|
| 148 |
-
font-size: 14px; color: var(--ink); font-weight: 500;
|
| 149 |
-
font-variant-numeric: tabular-nums;
|
| 150 |
-
}
|
| 151 |
-
|
| 152 |
-
/* ---- rollout trace: the observation stream the agent would receive ---- */
|
| 153 |
-
/* The trace lives on the left and the guess map on the right, so an
|
| 154 |
-
expanded map can never cover the observation stream. */
|
| 155 |
-
#trace {
|
| 156 |
-
position: absolute; top: 58px; left: 12px; width: 310px; z-index: 7;
|
| 157 |
-
max-height: calc(100% - 130px); display: flex; flex-direction: column;
|
| 158 |
-
background: var(--panel); border: 1px solid var(--edge); border-radius: 4px;
|
| 159 |
-
backdrop-filter: blur(8px); overflow: hidden;
|
| 160 |
-
}
|
| 161 |
-
#trace.hidden { display: none; }
|
| 162 |
-
#trace h4 {
|
| 163 |
-
margin: 0; padding: 8px 11px; font-size: 10.5px; font-weight: 500;
|
| 164 |
-
letter-spacing: .12em; text-transform: uppercase; color: var(--wire);
|
| 165 |
-
border-bottom: 1px solid var(--edge); display: flex;
|
| 166 |
-
justify-content: space-between;
|
| 167 |
-
}
|
| 168 |
-
#trace h4 em { color: var(--ink-faint); font-style: normal; letter-spacing: 0; }
|
| 169 |
-
#trace h4 > span:last-child { display: flex; align-items: center; gap: 8px; }
|
| 170 |
-
#collapse {
|
| 171 |
-
font-family: inherit; font-size: 13px; line-height: 1; cursor: pointer;
|
| 172 |
-
background: #232c33; border: 1px solid var(--edge); border-radius: 3px;
|
| 173 |
-
color: var(--ink); width: 22px; height: 19px; padding: 0;
|
| 174 |
-
}
|
| 175 |
-
#expand {
|
| 176 |
-
font-family: inherit; font-size: 11px; line-height: 1; cursor: pointer;
|
| 177 |
-
background: #232c33; border: 1px solid var(--edge); border-radius: 3px;
|
| 178 |
-
color: var(--ink); padding: 5px 9px;
|
| 179 |
-
}
|
| 180 |
-
#collapse:hover, #expand:hover { border-color: #5b6b78; color: #fff; }
|
| 181 |
-
/* With the trace hidden, a small control stays where it was. */
|
| 182 |
-
#expandWrap {
|
| 183 |
-
position: absolute; top: 58px; left: 12px; z-index: 7; display: none;
|
| 184 |
-
padding: 4px 5px;
|
| 185 |
-
}
|
| 186 |
-
#expandWrap.show { display: block; }
|
| 187 |
-
#steps { overflow-y: auto; padding: 4px 0; font-size: 11px; }
|
| 188 |
-
.step { padding: 6px 11px; border-bottom: 1px solid #1e262c; line-height: 1.5; }
|
| 189 |
-
.step:last-child { border-bottom: none; }
|
| 190 |
-
.step .op { color: var(--wire); }
|
| 191 |
-
.step .rw { float: right; color: var(--ink-faint); }
|
| 192 |
-
.step .fb { color: var(--ink-soft); display: block; margin-top: 2px; }
|
| 193 |
-
#agentview { border-top: 1px solid var(--edge); padding: 8px 11px; }
|
| 194 |
-
#agentview .cap {
|
| 195 |
-
font-size: 9.5px; letter-spacing: .1em; text-transform: uppercase;
|
| 196 |
-
color: var(--ink-faint); margin-bottom: 5px;
|
| 197 |
-
}
|
| 198 |
-
#agentview img { width: 100%; display: block; border-radius: 3px; }
|
| 199 |
-
|
| 200 |
-
/* ---- guess map ------------------------------------------------------- */
|
| 201 |
-
#mapwrap {
|
| 202 |
-
position: absolute; right: 12px; bottom: 12px; z-index: 6;
|
| 203 |
-
width: 300px; height: 210px; border: 1px solid var(--edge);
|
| 204 |
-
border-radius: 5px; overflow: hidden; opacity: .9; background: #191f24;
|
| 205 |
-
transition: width .24s ease, height .24s ease, opacity .24s ease,
|
| 206 |
-
right .24s ease, bottom .24s ease;
|
| 207 |
-
}
|
| 208 |
-
#mapwrap:hover, #mapwrap.big {
|
| 209 |
-
width: min(560px, 48vw); height: min(400px, 58vh); opacity: 1;
|
| 210 |
-
}
|
| 211 |
-
#mapwrap.reveal {
|
| 212 |
-
right: 50%; bottom: 50%; transform: translate(50%, 50%);
|
| 213 |
-
width: min(1000px, 86vw); height: min(600px, 72vh); opacity: 1;
|
| 214 |
-
}
|
| 215 |
-
/* The reveal is the whole point of the round, so it may cover the trace. */
|
| 216 |
-
#mapwrap.reveal { z-index: 13; }
|
| 217 |
-
#map { position: absolute; inset: 0; }
|
| 218 |
-
#submit {
|
| 219 |
-
position: absolute; left: 0; right: 0; bottom: 0; z-index: 3; width: 100%;
|
| 220 |
-
font-family: inherit; font-size: 12px; letter-spacing: .07em;
|
| 221 |
-
text-transform: uppercase; padding: 10px; border: none;
|
| 222 |
-
border-top: 1px solid var(--edge); background: #2b3540; color: #7e8b96;
|
| 223 |
-
cursor: not-allowed;
|
| 224 |
-
}
|
| 225 |
-
#submit.ready { background: var(--pin); color: #fff; cursor: pointer; }
|
| 226 |
-
#submit.ready:hover { background: #d64236; }
|
| 227 |
-
#mapwrap.reveal #submit { display: none; }
|
| 228 |
-
|
| 229 |
-
/* ---- result bar ------------------------------------------------------ */
|
| 230 |
-
#result {
|
| 231 |
-
position: absolute; left: 0; right: 0; bottom: 0; z-index: 12; display: none;
|
| 232 |
-
background: rgba(14, 18, 22, .96); border-top: 1px solid var(--edge);
|
| 233 |
-
padding: 13px 20px; backdrop-filter: blur(8px);
|
| 234 |
-
}
|
| 235 |
-
#result.show { display: block; }
|
| 236 |
-
#result .inner {
|
| 237 |
-
max-width: 1180px; margin: 0 auto; display: flex; align-items: center;
|
| 238 |
-
gap: 22px; flex-wrap: wrap;
|
| 239 |
-
}
|
| 240 |
-
#verdict { font-size: 17px; color: #fff; min-width: 140px; }
|
| 241 |
-
.stat { font-size: 10.5px; color: var(--ink-soft); }
|
| 242 |
-
.stat b {
|
| 243 |
-
display: block; font-size: 14px; color: var(--ink); font-weight: 500;
|
| 244 |
-
font-variant-numeric: tabular-nums; margin-top: 2px;
|
| 245 |
-
}
|
| 246 |
-
.stat.env b { color: var(--wire); }
|
| 247 |
-
#bar { flex: 1 1 140px; min-width: 100px; height: 6px; background: #232b32;
|
| 248 |
-
border-radius: 4px; overflow: hidden; }
|
| 249 |
-
#bar i { display: block; height: 100%; background: var(--good); width: 0; }
|
| 250 |
-
#advance {
|
| 251 |
-
font-family: inherit; font-size: 12px; text-transform: uppercase;
|
| 252 |
-
letter-spacing: .06em; padding: 10px 18px; background: #232b32;
|
| 253 |
-
color: var(--ink); border: 1px solid var(--edge); border-radius: 3px;
|
| 254 |
-
cursor: pointer;
|
| 255 |
-
}
|
| 256 |
-
#advance:hover { border-color: #465360; color: #fff; }
|
| 257 |
-
|
| 258 |
-
.dot {
|
| 259 |
-
display: inline-block; width: 7px; height: 7px; border-radius: 50%;
|
| 260 |
-
margin-right: 7px; background: var(--ink-faint);
|
| 261 |
-
vertical-align: 1px;
|
| 262 |
-
}
|
| 263 |
-
.dot.live { background: var(--good); }
|
| 264 |
-
.dot.bad { background: var(--pin); }
|
| 265 |
-
</style>
|
| 266 |
-
</head>
|
| 267 |
-
<body>
|
| 268 |
-
<div id="stage">
|
| 269 |
-
<div id="pano"></div>
|
| 270 |
-
|
| 271 |
-
<div class="hud" id="top">
|
| 272 |
-
<i id="conn" class="dot" title="environment session"></i><b>geoguesser_env</b>
|
| 273 |
-
<span>episode</span> <b class="ep" id="task">-</b>
|
| 274 |
-
<span>· frame</span> <b id="frame">-</b>
|
| 275 |
-
<span>· steps</span> <b id="stepsLeft">-</b>
|
| 276 |
-
</div>
|
| 277 |
-
<div class="hud" id="compass"><span>facing</span> <b id="heading">-</b>
|
| 278 |
-
<span>· fov</span> <b id="fov">-</b>
|
| 279 |
-
<span>· map</span> <b id="mapspan">-</b></div>
|
| 280 |
-
<div class="hud" id="score">
|
| 281 |
-
<div class="pts"><b id="total">0</b><span> pts</span></div>
|
| 282 |
-
<span id="played">0 episodes</span> · <span id="captured"></span>
|
| 283 |
-
</div>
|
| 284 |
-
<div class="hud" id="credit"></div>
|
| 285 |
-
<div class="hud" id="actions">
|
| 286 |
-
<div class="pad" id="padTurn">
|
| 287 |
-
<span class="padlabel">look</span>
|
| 288 |
-
<div class="btns">
|
| 289 |
-
<button id="turn-left" title="Turn 45° left and look (costs 0.01) — key: A or ←">
|
| 290 |
-
<span class="glyph">◀</span><span class="tag">left</span></button>
|
| 291 |
-
<button id="look-now" title="Look again where you are facing (costs 0.01) — key: L">
|
| 292 |
-
<span class="glyph">◎</span><span class="tag">look</span></button>
|
| 293 |
-
<button id="turn-right" title="Turn 45° right and look (costs 0.01) — key: D or →">
|
| 294 |
-
<span class="glyph">▶</span><span class="tag">right</span></button>
|
| 295 |
-
</div>
|
| 296 |
-
</div>
|
| 297 |
-
|
| 298 |
-
<div class="pad" id="padMove">
|
| 299 |
-
<span class="padlabel">walk</span>
|
| 300 |
-
<div class="btns">
|
| 301 |
-
<button id="move-fwd" title="Walk 15 m forward (costs 0.05) — key: W or ↑">
|
| 302 |
-
<span class="glyph">▲</span><span class="tag">forward</span></button>
|
| 303 |
-
<button id="move-back" title="Walk 15 m back (costs 0.05) — key: S or ↓">
|
| 304 |
-
<span class="glyph">▼</span><span class="tag">back</span></button>
|
| 305 |
-
</div>
|
| 306 |
-
</div>
|
| 307 |
-
|
| 308 |
-
<div class="pad" id="padZoom">
|
| 309 |
-
<span class="padlabel">zoom · <b id="fovValue">90°</b></span>
|
| 310 |
-
<div class="btns">
|
| 311 |
-
<button class="preset" data-fov="90" title="Wide view, 90° (costs 0.01)">wide</button>
|
| 312 |
-
<button class="preset" data-fov="50" title="Street level, 50° (costs 0.01)">street</button>
|
| 313 |
-
<button class="preset" data-fov="30" title="Read a sign, 30° (costs 0.01)">sign</button>
|
| 314 |
-
<input id="fovRange" type="range" min="20" max="110" step="5" value="90"
|
| 315 |
-
aria-label="field of view in degrees">
|
| 316 |
-
</div>
|
| 317 |
-
</div>
|
| 318 |
-
|
| 319 |
-
<div class="pad" id="padBudget">
|
| 320 |
-
<span class="padlabel">budget</span>
|
| 321 |
-
<div class="btns budget">
|
| 322 |
-
<b id="stepsBudget">12</b><span>left</span>
|
| 323 |
-
<span class="sep">·</span>
|
| 324 |
-
<b id="costBudget">0.00</b><span>spent</span>
|
| 325 |
-
</div>
|
| 326 |
-
</div>
|
| 327 |
-
</div>
|
| 328 |
-
|
| 329 |
-
<div class="hud" id="expandWrap">
|
| 330 |
-
<button id="expand" title="show the rollout trace (T)">+ trace</button>
|
| 331 |
-
</div>
|
| 332 |
-
|
| 333 |
-
<div id="trace">
|
| 334 |
-
<h4>
|
| 335 |
-
<span>rollout trace</span>
|
| 336 |
-
<span><em id="cost">cost 0.00</em>
|
| 337 |
-
<button id="collapse" title="hide the trace (T)">–</button></span>
|
| 338 |
-
</h4>
|
| 339 |
-
<div id="steps"></div>
|
| 340 |
-
<div id="agentview" style="display:none">
|
| 341 |
-
<div class="cap">what the agent sees</div>
|
| 342 |
-
<img id="agentimg" alt="the environment's own rendered observation">
|
| 343 |
-
</div>
|
| 344 |
-
</div>
|
| 345 |
-
|
| 346 |
-
<div id="mapwrap">
|
| 347 |
-
<div id="map"></div>
|
| 348 |
-
<button id="submit">click the map to place a pin</button>
|
| 349 |
-
</div>
|
| 350 |
-
|
| 351 |
-
<div id="result">
|
| 352 |
-
<div class="inner">
|
| 353 |
-
<div id="verdict">-</div>
|
| 354 |
-
<div class="stat">distance<b id="dist">-</b></div>
|
| 355 |
-
<div class="stat">points<b id="points">-</b></div>
|
| 356 |
-
<div class="stat env">env reward<b id="reward">-</b></div>
|
| 357 |
-
<div class="stat">score - cost<b id="breakdown2">-</b></div>
|
| 358 |
-
<div class="stat">true location<b id="truth">-</b></div>
|
| 359 |
-
<div id="bar"><i></i></div>
|
| 360 |
-
<button id="advance">load another episode</button>
|
| 361 |
-
</div>
|
| 362 |
-
</div>
|
| 363 |
-
|
| 364 |
-
</div>
|
| 365 |
-
|
| 366 |
-
<script src="__MAPLIBRE_JS__"></script>
|
| 367 |
-
<script src="__PANNELLUM_JS__"></script>
|
| 368 |
-
<script>
|
| 369 |
-
(function () {
|
| 370 |
-
"use strict";
|
| 371 |
-
const N_TASKS = __N_TASKS__;
|
| 372 |
-
const SPLIT = __SPLIT__;
|
| 373 |
-
const MAX_POINTS = __MAX_POINTS__;
|
| 374 |
-
// Every task-scoped request has to name its split, or index 12 of eval and
|
| 375 |
-
// index 12 of train are indistinguishable and the page reveals the wrong
|
| 376 |
-
// ground truth.
|
| 377 |
-
const q = (extra) => "?split=" + encodeURIComponent(SPLIT) + (extra || "");
|
| 378 |
-
const $ = (id) => document.getElementById(id);
|
| 379 |
-
|
| 380 |
-
let socket = null, ready = false, pending = null;
|
| 381 |
-
let viewer = null, map = null, guessMarker = null, truthMarker = null;
|
| 382 |
-
let lineAdded = false, guess = null, taskIndex = 0, compass = 0;
|
| 383 |
-
let taskMeta = null, frameIndex = 0;
|
| 384 |
-
let total = 0, played = 0, busy = false, cost = 0;
|
| 385 |
-
|
| 386 |
-
const DIRS = ["N", "NE", "E", "SE", "S", "SW", "W", "NW"];
|
| 387 |
-
const fmt = (la, lo) => la.toFixed(4) + ", " + lo.toFixed(4);
|
| 388 |
-
const yaw = () => (viewer ? ((viewer.getYaw() % 360) + 360) % 360 : 0);
|
| 389 |
-
const hfov = () => (viewer ? viewer.getHfov() : 90);
|
| 390 |
-
|
| 391 |
-
// ---- the environment, over the same WebSocket session API a client uses --
|
| 392 |
-
// Plain REST /step builds a fresh environment per request, so a stateful
|
| 393 |
-
// episode has to run over /ws. This page therefore plays exactly the
|
| 394 |
-
// rollout an agent would: one reset, a few charged steps, one terminal guess.
|
| 395 |
-
function connect() {
|
| 396 |
-
const scheme = location.protocol === "https:" ? "wss:" : "ws:";
|
| 397 |
-
socket = new WebSocket(scheme + "//" + location.host + "/ws");
|
| 398 |
-
socket.onopen = function () {
|
| 399 |
-
ready = true;
|
| 400 |
-
$("conn").className = "dot live";
|
| 401 |
-
$("conn").title = "environment session: connected";
|
| 402 |
-
startRound(requestedTask());
|
| 403 |
-
};
|
| 404 |
-
socket.onclose = function () {
|
| 405 |
-
ready = false;
|
| 406 |
-
$("conn").className = "dot bad";
|
| 407 |
-
$("conn").title = "environment session: disconnected";
|
| 408 |
-
};
|
| 409 |
-
socket.onerror = function () {
|
| 410 |
-
$("conn").className = "dot bad";
|
| 411 |
-
$("conn").title = "environment session: error";
|
| 412 |
-
};
|
| 413 |
-
socket.onmessage = function (event) {
|
| 414 |
-
const message = JSON.parse(event.data);
|
| 415 |
-
if (message.type === "error") {
|
| 416 |
-
addStep("error", message.data ? JSON.stringify(message.data) : "", null);
|
| 417 |
-
setBusy(false);
|
| 418 |
-
return;
|
| 419 |
-
}
|
| 420 |
-
if (message.type !== "observation") return;
|
| 421 |
-
// The wire format nests the observation and carries reward and done as
|
| 422 |
-
// siblings: {observation: {...}, reward, done, metadata}. Flatten it so
|
| 423 |
-
// callers read one object.
|
| 424 |
-
const payload = message.data || {};
|
| 425 |
-
const observation = Object.assign(
|
| 426 |
-
{}, payload.observation || payload,
|
| 427 |
-
{ reward: payload.reward, done: payload.done }
|
| 428 |
-
);
|
| 429 |
-
const handler = pending;
|
| 430 |
-
pending = null;
|
| 431 |
-
if (handler) handler(observation);
|
| 432 |
-
};
|
| 433 |
-
}
|
| 434 |
-
|
| 435 |
-
function send(type, data, handler) {
|
| 436 |
-
if (!ready) return;
|
| 437 |
-
pending = handler || null;
|
| 438 |
-
socket.send(JSON.stringify({ type: type, data: data || {} }));
|
| 439 |
-
}
|
| 440 |
-
|
| 441 |
-
// ---- trace ------------------------------------------------------------
|
| 442 |
-
function addStep(op, feedback, observation) {
|
| 443 |
-
const row = document.createElement("div");
|
| 444 |
-
row.className = "step";
|
| 445 |
-
let right = "";
|
| 446 |
-
if (observation && observation.steps_remaining !== undefined) {
|
| 447 |
-
right = "<span class='rw'>" + observation.steps_remaining + " left</span>";
|
| 448 |
-
}
|
| 449 |
-
row.innerHTML = "<span class='op'>" + op + "</span>" + right +
|
| 450 |
-
"<span class='fb'>" + (feedback || "") + "</span>";
|
| 451 |
-
$("steps").appendChild(row);
|
| 452 |
-
$("steps").scrollTop = $("steps").scrollHeight;
|
| 453 |
-
if (observation && observation.image_base64) {
|
| 454 |
-
const mime = observation.image_kind === "map" ? "png" : "jpeg";
|
| 455 |
-
$("agentimg").src = "data:image/" + mime + ";base64," + observation.image_base64;
|
| 456 |
-
$("agentview").style.display = "block";
|
| 457 |
-
}
|
| 458 |
-
}
|
| 459 |
-
|
| 460 |
-
function applyObservation(observation) {
|
| 461 |
-
if (!observation) return;
|
| 462 |
-
if (observation.steps_remaining !== undefined) {
|
| 463 |
-
$("stepsLeft").textContent = observation.steps_remaining;
|
| 464 |
-
$("stepsBudget").textContent = observation.steps_remaining;
|
| 465 |
-
}
|
| 466 |
-
if (observation.action_cost !== null && observation.action_cost !== undefined) {
|
| 467 |
-
cost = observation.action_cost;
|
| 468 |
-
$("cost").textContent = "cost " + cost.toFixed(2);
|
| 469 |
-
$("costBudget").textContent = cost.toFixed(2);
|
| 470 |
-
}
|
| 471 |
-
setControls(observation);
|
| 472 |
-
}
|
| 473 |
-
|
| 474 |
-
// ---- rounds -----------------------------------------------------------
|
| 475 |
-
function loadPano(index) {
|
| 476 |
-
fetch("/geoguesser/task/" + index + q())
|
| 477 |
-
.then((response) => response.json())
|
| 478 |
-
.then((meta) => {
|
| 479 |
-
taskMeta = meta;
|
| 480 |
-
frameIndex = meta.start_frame || 0;
|
| 481 |
-
const who = (meta.attribution || {}).creator_username;
|
| 482 |
-
$("credit").innerHTML =
|
| 483 |
-
"imagery © " + (who ? who : "Mapillary contributor") +
|
| 484 |
-
" via <a href='https://www.mapillary.com' target='_blank' rel='noopener'>Mapillary</a>" +
|
| 485 |
-
", <a href='https://creativecommons.org/licenses/by-sa/4.0/' target='_blank' rel='noopener'>CC BY-SA 4.0</a>";
|
| 486 |
-
showFrame(frameIndex, 0, 90);
|
| 487 |
-
});
|
| 488 |
-
}
|
| 489 |
-
|
| 490 |
-
/**
|
| 491 |
-
* Point the main viewer at one frame of the sequence.
|
| 492 |
-
*
|
| 493 |
-
* Called on reset and again after every move(), so walking forward actually
|
| 494 |
-
* changes what you are looking at rather than only what the trace shows.
|
| 495 |
-
* Heading and zoom carry over, because losing your orientation on every step
|
| 496 |
-
* would make navigation useless.
|
| 497 |
-
*/
|
| 498 |
-
function showFrame(index, keepYaw, keepHfov) {
|
| 499 |
-
frameIndex = index;
|
| 500 |
-
const frames = (taskMeta && taskMeta.frames) || [];
|
| 501 |
-
const frame = frames[index] || {};
|
| 502 |
-
compass = frame.compass_angle || 0;
|
| 503 |
-
if (frame.captured_at) {
|
| 504 |
-
$("captured").textContent = "captured " + frame.captured_at;
|
| 505 |
-
}
|
| 506 |
-
$("frame").textContent = index + "/" + Math.max(0, frames.length - 1);
|
| 507 |
-
if (viewer) { viewer.destroy(); viewer = null; }
|
| 508 |
-
viewer = pannellum.viewer("pano", {
|
| 509 |
-
type: "equirectangular",
|
| 510 |
-
panorama: "/geoguesser/pano/" + taskIndex + "/" + index + q(),
|
| 511 |
-
autoLoad: true, showControls: false, northOffset: compass,
|
| 512 |
-
yaw: keepYaw, hfov: keepHfov,
|
| 513 |
-
minHfov: 20, maxHfov: 110, compass: false, friction: 0.15,
|
| 514 |
-
});
|
| 515 |
-
viewer.on("mouseup", updateHud);
|
| 516 |
-
viewer.on("touchend", updateHud);
|
| 517 |
-
viewer.on("zoomchange", updateHud);
|
| 518 |
-
viewer.on("load", updateHud);
|
| 519 |
-
setTimeout(updateHud, 400);
|
| 520 |
-
}
|
| 521 |
-
|
| 522 |
-
function setControls(observation) {
|
| 523 |
-
if (!observation) return;
|
| 524 |
-
const tools = observation.available_tools || [];
|
| 525 |
-
const canLook = tools.indexOf("look") !== -1;
|
| 526 |
-
const canMove = tools.indexOf("move") !== -1;
|
| 527 |
-
$("padTurn").classList.toggle("gone", !canLook);
|
| 528 |
-
$("padZoom").classList.toggle("gone", !canLook);
|
| 529 |
-
$("look-now").classList.toggle("gone", !canLook);
|
| 530 |
-
$("padMove").classList.toggle(
|
| 531 |
-
"gone",
|
| 532 |
-
!canMove ||
|
| 533 |
-
(!observation.can_move_forward && !observation.can_move_backward)
|
| 534 |
-
);
|
| 535 |
-
$("move-fwd").classList.toggle("gone", !observation.can_move_forward);
|
| 536 |
-
$("move-back").classList.toggle("gone", !observation.can_move_backward);
|
| 537 |
-
if (observation.fov_deg) {
|
| 538 |
-
const fov = Math.round(observation.fov_deg);
|
| 539 |
-
$("fovRange").value = String(fov);
|
| 540 |
-
$("fovValue").textContent = fov + "\u00b0";
|
| 541 |
-
markPreset(fov);
|
| 542 |
-
}
|
| 543 |
-
}
|
| 544 |
-
|
| 545 |
-
function updateHud() {
|
| 546 |
-
if (!viewer) return;
|
| 547 |
-
const y = yaw();
|
| 548 |
-
$("heading").textContent = DIRS[Math.round(y / 45) % 8] + " " + y.toFixed(0) + "°";
|
| 549 |
-
$("fov").textContent = hfov().toFixed(0) + "°";
|
| 550 |
-
}
|
| 551 |
-
|
| 552 |
-
// ---- charged actions, executed by the environment ---------------------
|
| 553 |
-
function setBusy(value) {
|
| 554 |
-
busy = value;
|
| 555 |
-
$("actions").classList.toggle("working", value);
|
| 556 |
-
if (value) {
|
| 557 |
-
$("stepsBudget").textContent = "\u2026";
|
| 558 |
-
}
|
| 559 |
-
}
|
| 560 |
-
|
| 561 |
-
function step(op, data, label) {
|
| 562 |
-
if (busy || !ready) return;
|
| 563 |
-
setBusy(true);
|
| 564 |
-
send("step", Object.assign({ op: op }, data), function (observation) {
|
| 565 |
-
setBusy(false);
|
| 566 |
-
applyObservation(observation);
|
| 567 |
-
addStep(label, observation.feedback, observation);
|
| 568 |
-
const meta = observation.metadata || {};
|
| 569 |
-
if (op === "move" && meta.frame_index !== undefined &&
|
| 570 |
-
meta.frame_index !== frameIndex) {
|
| 571 |
-
// Keep the player facing the same way through the step.
|
| 572 |
-
showFrame(meta.frame_index, yaw(), hfov());
|
| 573 |
-
}
|
| 574 |
-
if (op === "look" && data && data.fov_deg && viewer) {
|
| 575 |
-
// The main view and the agent's view should never disagree.
|
| 576 |
-
viewer.setHfov(data.fov_deg);
|
| 577 |
-
if (data.heading_deg !== undefined) viewer.setYaw(data.heading_deg);
|
| 578 |
-
}
|
| 579 |
-
if (op === "guess") reveal(observation);
|
| 580 |
-
});
|
| 581 |
-
}
|
| 582 |
-
|
| 583 |
-
// The pad is the agent's action set, not a viewer control: every button is a
|
| 584 |
-
// charged environment step, and the panorama follows the result. Dragging the
|
| 585 |
-
// scene stays free, for orientation only.
|
| 586 |
-
function lookAt(heading, fov) {
|
| 587 |
-
const wrapped = ((Math.round(heading) % 360) + 360) % 360;
|
| 588 |
-
step(
|
| 589 |
-
"look",
|
| 590 |
-
{ heading_deg: wrapped, pitch_deg: 0, fov_deg: Math.round(fov) },
|
| 591 |
-
"look(heading=" + wrapped + ", fov=" + Math.round(fov) + ")"
|
| 592 |
-
);
|
| 593 |
-
}
|
| 594 |
-
|
| 595 |
-
$("turn-left").onclick = function () { lookAt(yaw() - 45, hfov()); };
|
| 596 |
-
$("turn-right").onclick = function () { lookAt(yaw() + 45, hfov()); };
|
| 597 |
-
$("look-now").onclick = function () { lookAt(yaw(), hfov()); };
|
| 598 |
-
|
| 599 |
-
Array.prototype.forEach.call(
|
| 600 |
-
document.querySelectorAll("#padZoom .preset"),
|
| 601 |
-
function (button) {
|
| 602 |
-
button.onclick = function () {
|
| 603 |
-
const fov = parseInt(button.dataset.fov, 10);
|
| 604 |
-
$("fovRange").value = String(fov);
|
| 605 |
-
$("fovValue").textContent = fov + "\u00b0";
|
| 606 |
-
markPreset(fov);
|
| 607 |
-
lookAt(yaw(), fov);
|
| 608 |
-
};
|
| 609 |
-
}
|
| 610 |
-
);
|
| 611 |
-
|
| 612 |
-
function markPreset(fov) {
|
| 613 |
-
Array.prototype.forEach.call(
|
| 614 |
-
document.querySelectorAll("#padZoom .preset"),
|
| 615 |
-
function (button) {
|
| 616 |
-
button.classList.toggle("on", parseInt(button.dataset.fov, 10) === fov);
|
| 617 |
-
}
|
| 618 |
-
);
|
| 619 |
-
}
|
| 620 |
-
$("move-fwd").onclick = function () {
|
| 621 |
-
step("move", { direction: "forward", meters: 15 }, "move(forward, 15m)");
|
| 622 |
-
};
|
| 623 |
-
$("move-back").onclick = function () {
|
| 624 |
-
step("move", { direction: "backward", meters: 15 }, "move(backward, 15m)");
|
| 625 |
-
};
|
| 626 |
-
|
| 627 |
-
// The slider reads out live but only spends a step on release, so dragging it
|
| 628 |
-
// does not burn the budget.
|
| 629 |
-
$("fovRange").addEventListener("input", function () {
|
| 630 |
-
const fov = parseInt($("fovRange").value, 10);
|
| 631 |
-
$("fovValue").textContent = fov + "\u00b0";
|
| 632 |
-
markPreset(fov);
|
| 633 |
-
});
|
| 634 |
-
$("fovRange").addEventListener("change", function () {
|
| 635 |
-
lookAt(yaw(), parseInt($("fovRange").value, 10));
|
| 636 |
-
});
|
| 637 |
-
function setTrace(visible) {
|
| 638 |
-
$("trace").classList.toggle("hidden", !visible);
|
| 639 |
-
$("expandWrap").classList.toggle("show", !visible);
|
| 640 |
-
}
|
| 641 |
-
$("collapse").onclick = function () { setTrace(false); };
|
| 642 |
-
$("expand").onclick = function () { setTrace(true); };
|
| 643 |
-
|
| 644 |
-
// The picker is the human equivalent of reset(task_index=k): the same call an
|
| 645 |
-
// eval harness makes, so a person can replay exactly the episode an agent saw.
|
| 646 |
-
/** Task index requested in the page URL, when the host supplied one. */
|
| 647 |
-
function requestedTask() {
|
| 648 |
-
const value = new URLSearchParams(location.search).get("task");
|
| 649 |
-
if (value === null || value === "" || value === "random") return undefined;
|
| 650 |
-
const parsed = parseInt(value, 10);
|
| 651 |
-
return Number.isFinite(parsed) ? parsed : undefined;
|
| 652 |
-
}
|
| 653 |
-
|
| 654 |
-
function startRound(index) {
|
| 655 |
-
cost = 0;
|
| 656 |
-
$("cost").textContent = "cost 0.00";
|
| 657 |
-
$("steps").innerHTML = "";
|
| 658 |
-
$("agentview").style.display = "none";
|
| 659 |
-
clearRound();
|
| 660 |
-
const wanted = index === undefined
|
| 661 |
-
? Math.floor(Math.random() * N_TASKS)
|
| 662 |
-
: ((index % N_TASKS) + N_TASKS) % N_TASKS;
|
| 663 |
-
taskIndex = wanted;
|
| 664 |
-
send("reset", { split: SPLIT, index: wanted }, function (observation) {
|
| 665 |
-
const meta = observation.metadata || {};
|
| 666 |
-
taskIndex = meta.task_index !== undefined ? meta.task_index : wanted;
|
| 667 |
-
$("task").textContent = taskIndex;
|
| 668 |
-
$("captured").textContent = "captured " + (observation.captured_at || "unknown");
|
| 669 |
-
applyObservation(observation);
|
| 670 |
-
addStep(
|
| 671 |
-
"reset(split='" + SPLIT + "', index=" + taskIndex + ")",
|
| 672 |
-
"episode started · " + (observation.available_tools || []).length +
|
| 673 |
-
" tools registered",
|
| 674 |
-
observation
|
| 675 |
-
);
|
| 676 |
-
loadPano(taskIndex);
|
| 677 |
-
});
|
| 678 |
-
}
|
| 679 |
-
|
| 680 |
-
// ---- map --------------------------------------------------------------
|
| 681 |
-
const WORLD = [[-179, -58], [179, 76]];
|
| 682 |
-
map = new maplibregl.Map({
|
| 683 |
-
container: "map", style: "__OPENFREEMAP_STYLE__",
|
| 684 |
-
center: [0, 12], zoom: 0, minZoom: -2,
|
| 685 |
-
attributionControl: { compact: true }, dragRotate: false,
|
| 686 |
-
// Without this the world repeats horizontally, which reads as a rendering
|
| 687 |
-
// bug at the zoom levels a small guess map uses.
|
| 688 |
-
renderWorldCopies: false,
|
| 689 |
-
});
|
| 690 |
-
map.on("load", () => map.fitBounds(WORLD, { padding: 6, duration: 0 }));
|
| 691 |
-
// Exposed so the page can be driven from a test harness or the console.
|
| 692 |
-
window.__ggMap = map;
|
| 693 |
-
window.__ggState = function () {
|
| 694 |
-
return {
|
| 695 |
-
busy: busy,
|
| 696 |
-
ready: ready,
|
| 697 |
-
pendingHandler: !!pending,
|
| 698 |
-
socket: socket ? socket.readyState : null,
|
| 699 |
-
steps: document.querySelectorAll('#steps .step').length,
|
| 700 |
-
};
|
| 701 |
-
};
|
| 702 |
-
|
| 703 |
-
/**
|
| 704 |
-
* Half-width of the visible map, in degrees.
|
| 705 |
-
*
|
| 706 |
-
* The environment renders its own map from this, so a pin dropped while
|
| 707 |
-
* zoomed into a city comes back as a street-level map rather than a
|
| 708 |
-
* continental one. Without it the agent's view and the player's would
|
| 709 |
-
* disagree about how precisely the pin could be aimed.
|
| 710 |
-
*/
|
| 711 |
-
function currentSpanDeg() {
|
| 712 |
-
const bounds = map.getBounds();
|
| 713 |
-
const span = Math.abs(bounds.getEast() - bounds.getWest()) / 2;
|
| 714 |
-
return Math.min(180, Math.max(0.03, span));
|
| 715 |
-
}
|
| 716 |
-
|
| 717 |
-
function updateMapSpan() {
|
| 718 |
-
const span = currentSpanDeg();
|
| 719 |
-
$("mapspan").textContent =
|
| 720 |
-
span >= 1 ? span.toFixed(0) + "\u00b0" : (span * 111).toFixed(0) + " km";
|
| 721 |
-
}
|
| 722 |
-
map.on("zoomend", updateMapSpan);
|
| 723 |
-
map.on("moveend", updateMapSpan);
|
| 724 |
-
map.on("load", updateMapSpan);
|
| 725 |
-
|
| 726 |
-
map.on("click", function (event) {
|
| 727 |
-
if (busy || $("result").classList.contains("show")) return;
|
| 728 |
-
// Once you have committed to a pin the map stays open; letting it collapse
|
| 729 |
-
// on mouse-out makes it easy to lose the guess you were adjusting.
|
| 730 |
-
$("mapwrap").classList.add("big");
|
| 731 |
-
map.resize();
|
| 732 |
-
guess = event.lngLat;
|
| 733 |
-
if (guessMarker) guessMarker.remove();
|
| 734 |
-
guessMarker = new maplibregl.Marker({ color: "#c4332a" })
|
| 735 |
-
.setLngLat(guess).addTo(map);
|
| 736 |
-
$("submit").className = "ready";
|
| 737 |
-
$("submit").textContent = "submit guess";
|
| 738 |
-
// A pin is a real, charged environment step, so the map the agent would see
|
| 739 |
-
// comes back in the trace panel — framed at the zoom you are looking at, so
|
| 740 |
-
// the two views agree about how precisely the pin was aimed.
|
| 741 |
-
const span = currentSpanDeg();
|
| 742 |
-
step("pin", { lat: guess.lat, lon: guess.lng, span_deg: span },
|
| 743 |
-
"place_pin(" + guess.lat.toFixed(2) + ", " + guess.lng.toFixed(2) +
|
| 744 |
-
", span=" + span.toFixed(2) + ")");
|
| 745 |
-
});
|
| 746 |
-
|
| 747 |
-
$("submit").onclick = function () {
|
| 748 |
-
if (!guess) return;
|
| 749 |
-
step("guess", { lat: guess.lat, lon: guess.lng },
|
| 750 |
-
"submit_guess(" + guess.lat.toFixed(2) + ", " + guess.lng.toFixed(2) + ")");
|
| 751 |
-
};
|
| 752 |
-
|
| 753 |
-
function reveal(observation) {
|
| 754 |
-
if (observation.distance_km === null || observation.distance_km === undefined) {
|
| 755 |
-
addStep("guess rejected",
|
| 756 |
-
observation.feedback || "the environment returned no distance",
|
| 757 |
-
observation);
|
| 758 |
-
return;
|
| 759 |
-
}
|
| 760 |
-
const km = observation.distance_km;
|
| 761 |
-
const reward = observation.reward === null ? 0 : observation.reward;
|
| 762 |
-
const score = observation.score === null ? 0 : observation.score;
|
| 763 |
-
const points = Math.round(score * MAX_POINTS);
|
| 764 |
-
const truthLat = observation.true_lat, truthLon = observation.true_lon;
|
| 765 |
-
total += points;
|
| 766 |
-
played += 1;
|
| 767 |
-
$("played").textContent = played + (played === 1 ? " episode" : " episodes");
|
| 768 |
-
|
| 769 |
-
$("verdict").textContent =
|
| 770 |
-
km < 0.025 ? "Perfect." : km < 25 ? "Pinpoint." : km < 200 ? "Close." :
|
| 771 |
-
km < 1500 ? "Right region." : "Wrong continent.";
|
| 772 |
-
$("dist").textContent = km < 10 ? (km * 1000).toFixed(0) + " m" : km.toFixed(0) + " km";
|
| 773 |
-
$("points").textContent = points + " / " + MAX_POINTS;
|
| 774 |
-
$("reward").textContent = reward.toFixed(3);
|
| 775 |
-
$("breakdown2").textContent =
|
| 776 |
-
score.toFixed(3) + " - " + (observation.action_cost || 0).toFixed(2);
|
| 777 |
-
$("truth").textContent = fmt(truthLat, truthLon);
|
| 778 |
-
$("bar").firstElementChild.style.width = (reward * 100).toFixed(1) + "%";
|
| 779 |
-
$("total").textContent = total;
|
| 780 |
-
$("result").classList.add("show");
|
| 781 |
-
document.body.classList.add("revealing");
|
| 782 |
-
$("actions").style.opacity = "0";
|
| 783 |
-
|
| 784 |
-
truthMarker = new maplibregl.Marker({ color: "#5aa06e" })
|
| 785 |
-
.setLngLat([truthLon, truthLat]).addTo(map);
|
| 786 |
-
const line = {
|
| 787 |
-
type: "Feature",
|
| 788 |
-
geometry: {
|
| 789 |
-
type: "LineString",
|
| 790 |
-
coordinates: [[guess.lng, guess.lat], [truthLon, truthLat]],
|
| 791 |
-
},
|
| 792 |
-
};
|
| 793 |
-
if (lineAdded) {
|
| 794 |
-
map.getSource("shot").setData(line);
|
| 795 |
-
} else {
|
| 796 |
-
map.addSource("shot", { type: "geojson", data: line });
|
| 797 |
-
map.addLayer({
|
| 798 |
-
id: "shot", type: "line", source: "shot",
|
| 799 |
-
paint: { "line-color": "#c4332a", "line-width": 2.5, "line-dasharray": [2, 1.6] },
|
| 800 |
-
});
|
| 801 |
-
lineAdded = true;
|
| 802 |
-
}
|
| 803 |
-
$("mapwrap").classList.add("reveal");
|
| 804 |
-
map.resize();
|
| 805 |
-
setTimeout(function () {
|
| 806 |
-
map.fitBounds(
|
| 807 |
-
[[Math.min(guess.lng, truthLon), Math.min(guess.lat, truthLat)],
|
| 808 |
-
[Math.max(guess.lng, truthLon), Math.max(guess.lat, truthLat)]],
|
| 809 |
-
{ padding: 80, maxZoom: 7, duration: 900 }
|
| 810 |
-
);
|
| 811 |
-
}, 260);
|
| 812 |
-
}
|
| 813 |
-
|
| 814 |
-
function clearRound() {
|
| 815 |
-
guess = null;
|
| 816 |
-
setBusy(false);
|
| 817 |
-
if (guessMarker) { guessMarker.remove(); guessMarker = null; }
|
| 818 |
-
if (truthMarker) { truthMarker.remove(); truthMarker = null; }
|
| 819 |
-
if (lineAdded) {
|
| 820 |
-
map.getSource("shot").setData({
|
| 821 |
-
type: "Feature", geometry: { type: "LineString", coordinates: [] },
|
| 822 |
-
});
|
| 823 |
-
}
|
| 824 |
-
$("mapwrap").classList.remove("reveal", "big");
|
| 825 |
-
map.resize();
|
| 826 |
-
map.fitBounds(WORLD, { padding: 6, duration: 700 });
|
| 827 |
-
$("submit").className = "";
|
| 828 |
-
$("submit").textContent = "click the map to place a pin";
|
| 829 |
-
$("result").classList.remove("show");
|
| 830 |
-
document.body.classList.remove("revealing");
|
| 831 |
-
$("actions").style.opacity = "1";
|
| 832 |
-
}
|
| 833 |
-
|
| 834 |
-
// An episode is one guess, so there is no round to advance: the terminal
|
| 835 |
-
// control simply starts another episode.
|
| 836 |
-
$("advance").onclick = function () {
|
| 837 |
-
clearRound();
|
| 838 |
-
startRound();
|
| 839 |
-
};
|
| 840 |
-
|
| 841 |
-
document.addEventListener("keydown", function (event) {
|
| 842 |
-
if (event.key === "m" || event.key === "M") {
|
| 843 |
-
$("mapwrap").classList.toggle("big");
|
| 844 |
-
map.resize();
|
| 845 |
-
if (!guess && !busy) map.fitBounds(WORLD, { padding: 6, duration: 250 });
|
| 846 |
-
} else if (event.key === "t" || event.key === "T") {
|
| 847 |
-
setTrace($("trace").classList.contains("hidden"));
|
| 848 |
-
} else if (event.key === "ArrowUp" || event.key === "w") {
|
| 849 |
-
if (!$("move-fwd").classList.contains("gone")) $("move-fwd").click();
|
| 850 |
-
} else if (event.key === "ArrowDown" || event.key === "s") {
|
| 851 |
-
if (!$("move-back").classList.contains("gone")) $("move-back").click();
|
| 852 |
-
} else if (event.key === "l" || event.key === "L") {
|
| 853 |
-
if (!$("padTurn").classList.contains("gone")) $("look-now").click();
|
| 854 |
-
} else if (event.key === "ArrowLeft" || event.key === "a") {
|
| 855 |
-
if (!$("padTurn").classList.contains("gone")) $("turn-left").click();
|
| 856 |
-
} else if (event.key === "ArrowRight" || event.key === "d") {
|
| 857 |
-
if (!$("padTurn").classList.contains("gone")) $("turn-right").click();
|
| 858 |
-
} else if (event.key === "Enter") {
|
| 859 |
-
if ($("result").classList.contains("show")) { $("advance").click(); }
|
| 860 |
-
else { $("submit").click(); }
|
| 861 |
-
}
|
| 862 |
-
});
|
| 863 |
-
|
| 864 |
-
connect();
|
| 865 |
-
})();
|
| 866 |
-
</script>
|
| 867 |
-
</body>
|
| 868 |
-
</html>
|
| 869 |
-
"""
|
| 870 |
-
|
| 871 |
-
|
| 872 |
-
def play_page_html(splits: list[dict] | int, split: str | None = None) -> str:
|
| 873 |
-
"""
|
| 874 |
-
Return the standalone play page.
|
| 875 |
-
|
| 876 |
-
Served at `/geoguesser/play` and embedded in the Gradio tab through an
|
| 877 |
-
iframe. It is a full document rather than a fragment because `gr.HTML`
|
| 878 |
-
inserts markup without running `<script>` tags.
|
| 879 |
-
|
| 880 |
-
Args:
|
| 881 |
-
splits (`list[dict]` or `int`):
|
| 882 |
-
Split descriptors from [`~GeoGuesserEnvironment.list_splits`]. A
|
| 883 |
-
bare integer is accepted as a task count for callers predating
|
| 884 |
-
splits.
|
| 885 |
-
split (`str`, *optional*):
|
| 886 |
-
Which split the page should play. Defaults to the split marked
|
| 887 |
-
`default`, else the first one.
|
| 888 |
-
|
| 889 |
-
Returns:
|
| 890 |
-
`str`: A complete HTML document.
|
| 891 |
-
"""
|
| 892 |
-
if isinstance(splits, int):
|
| 893 |
-
descriptors = [
|
| 894 |
-
{"name": "train", "num_tasks": splits, "default": True, "type": "train"}
|
| 895 |
-
]
|
| 896 |
-
else:
|
| 897 |
-
descriptors = list(splits) or [
|
| 898 |
-
{"name": "train", "num_tasks": 1, "default": True, "type": "train"}
|
| 899 |
-
]
|
| 900 |
-
chosen = next(
|
| 901 |
-
(d for d in descriptors if d["name"] == split),
|
| 902 |
-
next((d for d in descriptors if d.get("default")), descriptors[0]),
|
| 903 |
-
)
|
| 904 |
-
replacements = {
|
| 905 |
-
"__SPLIT__": json.dumps(chosen["name"]),
|
| 906 |
-
"__N_TASKS__": str(max(1, int(chosen.get("num_tasks", 1)))),
|
| 907 |
-
"__MAX_POINTS__": str(MAX_POINTS_PER_ROUND),
|
| 908 |
-
"__MAPLIBRE_JS__": MAPLIBRE_JS,
|
| 909 |
-
"__MAPLIBRE_CSS__": MAPLIBRE_CSS,
|
| 910 |
-
"__PANNELLUM_JS__": PANNELLUM_JS,
|
| 911 |
-
"__PANNELLUM_CSS__": PANNELLUM_CSS,
|
| 912 |
-
"__OPENFREEMAP_STYLE__": OPENFREEMAP_STYLE,
|
| 913 |
-
}
|
| 914 |
-
page = _TEMPLATE
|
| 915 |
-
for token, value in replacements.items():
|
| 916 |
-
page = page.replace(token, value)
|
| 917 |
-
return page
|
| 918 |
-
|
| 919 |
-
|
| 920 |
-
def _iframe(task: str | int = "random", split: str = "") -> str:
|
| 921 |
-
"""Markup for the play iframe, pointed at one task of one split.
|
| 922 |
-
|
| 923 |
-
Args:
|
| 924 |
-
task (`str` or `int`, *optional*, defaults to `"random"`):
|
| 925 |
-
Task index to open, or `"random"`.
|
| 926 |
-
split (`str`, *optional*):
|
| 927 |
-
Split to play. Empty means the server's default split.
|
| 928 |
-
|
| 929 |
-
Returns:
|
| 930 |
-
`str`: An iframe element. Re-rendering it with a different task is what
|
| 931 |
-
makes the Gradio controls reload the round, since the page reads its
|
| 932 |
-
task from the URL.
|
| 933 |
-
"""
|
| 934 |
-
query = f"?task={task}"
|
| 935 |
-
if split:
|
| 936 |
-
query += f"&split={urllib.parse.quote(split)}"
|
| 937 |
-
return (
|
| 938 |
-
f'<iframe src="/geoguesser/play{query}" '
|
| 939 |
-
'style="width:100%;height:720px;border:1px solid #2c353d;'
|
| 940 |
-
'border-radius:6px" allow="fullscreen"></iframe>'
|
| 941 |
-
)
|
| 942 |
-
|
| 943 |
-
|
| 944 |
-
def build_geoguesser_gradio_app(
|
| 945 |
-
web_manager: Any,
|
| 946 |
-
action_fields: List[Dict[str, Any]],
|
| 947 |
-
metadata: Optional[Any],
|
| 948 |
-
is_chat_env: bool,
|
| 949 |
-
title: str,
|
| 950 |
-
quick_start_md: str,
|
| 951 |
-
) -> gr.Blocks:
|
| 952 |
-
"""
|
| 953 |
-
Build the human-play tab.
|
| 954 |
-
|
| 955 |
-
The episode controls live here, on the Gradio side, rather than inside the
|
| 956 |
-
page: picking a task is orchestration, the same `reset(task_index=k)` an
|
| 957 |
-
eval harness calls, so it belongs with the host controls and not among the
|
| 958 |
-
in-game HUD.
|
| 959 |
-
|
| 960 |
-
Args:
|
| 961 |
-
web_manager (`Any`):
|
| 962 |
-
The playground's environment manager, unused here.
|
| 963 |
-
action_fields (`list[dict]`):
|
| 964 |
-
Action schema fields, unused here.
|
| 965 |
-
metadata (`Any`, *optional*):
|
| 966 |
-
Environment metadata, unused here.
|
| 967 |
-
is_chat_env (`bool`):
|
| 968 |
-
Whether the env is chat-shaped, unused here.
|
| 969 |
-
title (`str`):
|
| 970 |
-
Playground title.
|
| 971 |
-
quick_start_md (`str`):
|
| 972 |
-
Quick-start markdown, unused here.
|
| 973 |
-
|
| 974 |
-
Returns:
|
| 975 |
-
`gradio.Blocks`: The play tab, hosting `/geoguesser/play` in an iframe.
|
| 976 |
-
"""
|
| 977 |
-
# Ask the server which splits it actually serves, rather than re-deriving
|
| 978 |
-
# them here from environment variables and drifting out of step with it.
|
| 979 |
-
descriptors: list[dict] = []
|
| 980 |
-
try:
|
| 981 |
-
from .app import ACTIVE_DEFAULT_SPLIT, create_geoguesser_environment
|
| 982 |
-
|
| 983 |
-
descriptors = create_geoguesser_environment().list_splits()
|
| 984 |
-
default_split = ACTIVE_DEFAULT_SPLIT
|
| 985 |
-
except Exception: # pragma: no cover - the page still works without counts
|
| 986 |
-
default_split = "train"
|
| 987 |
-
if not descriptors:
|
| 988 |
-
descriptors = [
|
| 989 |
-
{"name": default_split, "num_tasks": 1, "default": True, "type": "train"}
|
| 990 |
-
]
|
| 991 |
-
|
| 992 |
-
counts = {d["name"]: max(1, int(d.get("num_tasks", 1))) for d in descriptors}
|
| 993 |
-
names = list(counts)
|
| 994 |
-
if default_split not in counts:
|
| 995 |
-
default_split = names[0]
|
| 996 |
-
|
| 997 |
-
def _label(split: str) -> str:
|
| 998 |
-
return f"reset(index=) · 0 to {counts[split] - 1}"
|
| 999 |
-
|
| 1000 |
-
with gr.Blocks(title="Geoguesser Environment") as blocks:
|
| 1001 |
-
with gr.Row():
|
| 1002 |
-
split_box = gr.Dropdown(
|
| 1003 |
-
choices=names,
|
| 1004 |
-
value=default_split,
|
| 1005 |
-
label="reset(split=)",
|
| 1006 |
-
scale=1,
|
| 1007 |
-
interactive=len(names) > 1,
|
| 1008 |
-
)
|
| 1009 |
-
task_box = gr.Number(
|
| 1010 |
-
value=0,
|
| 1011 |
-
minimum=0,
|
| 1012 |
-
maximum=counts[default_split] - 1,
|
| 1013 |
-
step=1,
|
| 1014 |
-
precision=0,
|
| 1015 |
-
label=_label(default_split),
|
| 1016 |
-
scale=2,
|
| 1017 |
-
)
|
| 1018 |
-
load_button = gr.Button("load episode", variant="primary", scale=1)
|
| 1019 |
-
random_button = gr.Button("random episode", scale=1)
|
| 1020 |
-
frame = gr.HTML(value=_iframe("random", default_split), show_label=False)
|
| 1021 |
-
|
| 1022 |
-
def _on_split(split: str):
|
| 1023 |
-
"""Re-range the index box so it cannot address a missing task."""
|
| 1024 |
-
split = split or default_split
|
| 1025 |
-
return gr.update(maximum=counts[split] - 1, value=0, label=_label(split))
|
| 1026 |
-
|
| 1027 |
-
split_box.change(fn=_on_split, inputs=split_box, outputs=task_box)
|
| 1028 |
-
load_button.click(
|
| 1029 |
-
fn=lambda index, split: _iframe(int(index or 0), split or default_split),
|
| 1030 |
-
inputs=[task_box, split_box],
|
| 1031 |
-
outputs=frame,
|
| 1032 |
-
)
|
| 1033 |
-
random_button.click(
|
| 1034 |
-
fn=lambda split: _iframe(
|
| 1035 |
-
random.randrange(counts[split or default_split]),
|
| 1036 |
-
split or default_split,
|
| 1037 |
-
),
|
| 1038 |
-
inputs=split_box,
|
| 1039 |
-
outputs=frame,
|
| 1040 |
-
)
|
| 1041 |
-
return blocks
|
| 1042 |
-
|
| 1043 |
-
|
| 1044 |
-
__all__ = ["build_geoguesser_gradio_app", "play_page_html"]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/parser.py
DELETED
|
@@ -1,149 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Extract coordinates from a model's free-text reply.
|
| 4 |
-
|
| 5 |
-
Models emit reasoning and coordinates together, in many shapes. The parser
|
| 6 |
-
accepts what they actually produce rather than demanding a schema, and returns
|
| 7 |
-
`None` when nothing usable is present so the failure lands in the reward
|
| 8 |
-
instead of raising.
|
| 9 |
-
"""
|
| 10 |
-
|
| 11 |
-
from __future__ import annotations
|
| 12 |
-
|
| 13 |
-
import re
|
| 14 |
-
from dataclasses import dataclass
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
# 48.8584, 2.2945 | -16.49 / -68.12 | lat: 12.9 lon: 77.5
|
| 18 |
-
_DECIMAL_PAIR = re.compile(
|
| 19 |
-
r"(-?\d{1,3}(?:\.\d+)?)\s*(?:,|/|;|\s+and\s+|\s+)\s*(-?\d{1,3}(?:\.\d+)?)"
|
| 20 |
-
)
|
| 21 |
-
|
| 22 |
-
# 48°51'29"N 2°17'40"E
|
| 23 |
-
_DMS = re.compile(
|
| 24 |
-
r"(\d{1,3})\s*[°d]\s*(\d{1,2})?\s*['′m]?\s*(\d{1,2}(?:\.\d+)?)?"
|
| 25 |
-
r"\s*[\"″s]?\s*([NSEW])",
|
| 26 |
-
re.IGNORECASE,
|
| 27 |
-
)
|
| 28 |
-
|
| 29 |
-
_LABELLED = re.compile(
|
| 30 |
-
r"lat(?:itude)?\s*[:=]\s*(-?\d{1,3}(?:\.\d+)?)"
|
| 31 |
-
r".{0,40}?"
|
| 32 |
-
r"lon(?:g|gitude)?\s*[:=]\s*(-?\d{1,3}(?:\.\d+)?)",
|
| 33 |
-
re.IGNORECASE | re.DOTALL,
|
| 34 |
-
)
|
| 35 |
-
|
| 36 |
-
_TAG = re.compile(r"<guess>(.*?)</guess>", re.IGNORECASE | re.DOTALL)
|
| 37 |
-
|
| 38 |
-
_JSON_ISH = re.compile(
|
| 39 |
-
r"\"lat(?:itude)?\"\s*:\s*(-?\d{1,3}(?:\.\d+)?)"
|
| 40 |
-
r".{0,60}?"
|
| 41 |
-
r"\"lon(?:g|gitude)?\"\s*:\s*(-?\d{1,3}(?:\.\d+)?)",
|
| 42 |
-
re.IGNORECASE | re.DOTALL,
|
| 43 |
-
)
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
@dataclass
|
| 47 |
-
class ParsedGuess:
|
| 48 |
-
"""Outcome of parsing a reply.
|
| 49 |
-
|
| 50 |
-
Attributes:
|
| 51 |
-
lat (`float` or `None`):
|
| 52 |
-
Latitude, or `None` when nothing could be extracted.
|
| 53 |
-
lon (`float` or `None`):
|
| 54 |
-
Longitude, or `None` when nothing could be extracted.
|
| 55 |
-
source (`str`):
|
| 56 |
-
Which pattern matched: `"tag"`, `"json"`, `"labelled"`, `"dms"`,
|
| 57 |
-
`"decimal"` or `"none"`.
|
| 58 |
-
note (`str`):
|
| 59 |
-
Short explanation, safe to show the model as feedback.
|
| 60 |
-
"""
|
| 61 |
-
|
| 62 |
-
lat: float | None
|
| 63 |
-
lon: float | None
|
| 64 |
-
source: str
|
| 65 |
-
note: str = ""
|
| 66 |
-
|
| 67 |
-
@property
|
| 68 |
-
def ok(self) -> bool:
|
| 69 |
-
"""Whether a usable coordinate pair was extracted."""
|
| 70 |
-
return self.lat is not None and self.lon is not None
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
def _valid(lat: float, lon: float) -> bool:
|
| 74 |
-
return -90.0 <= lat <= 90.0 and -180.0 <= lon <= 180.0
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
def _dms_to_decimal(deg: str, minute: str | None, sec: str | None, hemi: str) -> float:
|
| 78 |
-
value = float(deg) + float(minute or 0) / 60 + float(sec or 0) / 3600
|
| 79 |
-
return -value if hemi.upper() in ("S", "W") else value
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
def parse_guess(response: str) -> ParsedGuess:
|
| 83 |
-
"""
|
| 84 |
-
Pull a coordinate pair out of a model reply.
|
| 85 |
-
|
| 86 |
-
Patterns are tried most explicit first, so a `<guess>` tag or a labelled
|
| 87 |
-
`lat:`/`lon:` pair wins over a bare number pair that might be a date or a
|
| 88 |
-
step count.
|
| 89 |
-
|
| 90 |
-
Args:
|
| 91 |
-
response (`str`):
|
| 92 |
-
The model's unedited reply.
|
| 93 |
-
|
| 94 |
-
Returns:
|
| 95 |
-
[`ParsedGuess`]: The extracted coordinates, or a result whose `ok` is
|
| 96 |
-
`False` with a `note` explaining what was wrong.
|
| 97 |
-
|
| 98 |
-
Examples:
|
| 99 |
-
|
| 100 |
-
```python
|
| 101 |
-
parse_guess("I think coastal Portugal. <guess>38.72, -9.14</guess>")
|
| 102 |
-
```
|
| 103 |
-
"""
|
| 104 |
-
if not response or not response.strip():
|
| 105 |
-
return ParsedGuess(None, None, "none", "Empty response.")
|
| 106 |
-
|
| 107 |
-
tagged = _TAG.search(response)
|
| 108 |
-
haystacks = [(tagged.group(1), "tag")] if tagged else []
|
| 109 |
-
haystacks.append((response, "body"))
|
| 110 |
-
|
| 111 |
-
for text, origin in haystacks:
|
| 112 |
-
for pattern, name in ((_JSON_ISH, "json"), (_LABELLED, "labelled")):
|
| 113 |
-
m = pattern.search(text)
|
| 114 |
-
if m:
|
| 115 |
-
lat, lon = float(m.group(1)), float(m.group(2))
|
| 116 |
-
if _valid(lat, lon):
|
| 117 |
-
src = name if origin == "body" else "tag"
|
| 118 |
-
return ParsedGuess(lat, lon, src)
|
| 119 |
-
return ParsedGuess(
|
| 120 |
-
None, None, "none", f"Coordinates out of range: {lat}, {lon}."
|
| 121 |
-
)
|
| 122 |
-
|
| 123 |
-
dms = _DMS.findall(text)
|
| 124 |
-
if len(dms) >= 2:
|
| 125 |
-
lat_m = next((d for d in dms if d[3].upper() in ("N", "S")), None)
|
| 126 |
-
lon_m = next((d for d in dms if d[3].upper() in ("E", "W")), None)
|
| 127 |
-
if lat_m and lon_m:
|
| 128 |
-
lat = _dms_to_decimal(*lat_m)
|
| 129 |
-
lon = _dms_to_decimal(*lon_m)
|
| 130 |
-
if _valid(lat, lon):
|
| 131 |
-
return ParsedGuess(lat, lon, "dms")
|
| 132 |
-
|
| 133 |
-
m = _DECIMAL_PAIR.search(text)
|
| 134 |
-
if m:
|
| 135 |
-
lat, lon = float(m.group(1)), float(m.group(2))
|
| 136 |
-
if _valid(lat, lon):
|
| 137 |
-
src = "decimal" if origin == "body" else "tag"
|
| 138 |
-
return ParsedGuess(lat, lon, src)
|
| 139 |
-
return ParsedGuess(
|
| 140 |
-
None, None, "none", f"Coordinates out of range: {lat}, {lon}."
|
| 141 |
-
)
|
| 142 |
-
|
| 143 |
-
return ParsedGuess(
|
| 144 |
-
None,
|
| 145 |
-
None,
|
| 146 |
-
"none",
|
| 147 |
-
"No coordinates found. Reply with a latitude and longitude, for "
|
| 148 |
-
"example <guess>48.8584, 2.2945</guess>.",
|
| 149 |
-
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/render/__init__.py
DELETED
|
@@ -1,3 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Image rendering for the GeoGuesser environment."""
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/render/minimap.py
DELETED
|
@@ -1,824 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""The guess map, rendered offline.
|
| 4 |
-
|
| 5 |
-
Everything here runs against bundled Natural Earth vectors, so a pin costs no
|
| 6 |
-
network call and always renders the same bytes. Map detail is a function of
|
| 7 |
-
zoom alone and never of proximity to the target: prefetching finer data around
|
| 8 |
-
task locations would turn the map into a ground-truth oracle.
|
| 9 |
-
|
| 10 |
-
The renderer answers only "where did the agent point?". It never draws, names
|
| 11 |
-
or hints at the true location.
|
| 12 |
-
"""
|
| 13 |
-
|
| 14 |
-
from __future__ import annotations
|
| 15 |
-
|
| 16 |
-
import functools
|
| 17 |
-
import hashlib
|
| 18 |
-
import io
|
| 19 |
-
import json
|
| 20 |
-
import logging
|
| 21 |
-
import math
|
| 22 |
-
import os
|
| 23 |
-
import pathlib
|
| 24 |
-
import urllib.parse
|
| 25 |
-
import urllib.request
|
| 26 |
-
from dataclasses import dataclass
|
| 27 |
-
|
| 28 |
-
import matplotlib
|
| 29 |
-
|
| 30 |
-
matplotlib.use("Agg")
|
| 31 |
-
|
| 32 |
-
import matplotlib.patheffects as path_effects # noqa: E402
|
| 33 |
-
import matplotlib.pyplot as plt # noqa: E402
|
| 34 |
-
from matplotlib.collections import LineCollection # noqa: E402
|
| 35 |
-
from matplotlib.patches import Polygon as MplPolygon # noqa: E402
|
| 36 |
-
from matplotlib.path import Path as MplPath # noqa: E402
|
| 37 |
-
from PIL import Image # noqa: E402
|
| 38 |
-
|
| 39 |
-
from ..scoring import haversine_km as _haversine_km # noqa: E402
|
| 40 |
-
|
| 41 |
-
logger = logging.getLogger(__name__)
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
GEO_DIR = pathlib.Path(__file__).resolve().parents[2] / "data" / "geo"
|
| 45 |
-
COUNTRIES_FILE = GEO_DIR / "ne_110m_admin_0_countries.geojson"
|
| 46 |
-
CITIES_FILE = GEO_DIR / "ne_50m_populated_places.geojson"
|
| 47 |
-
|
| 48 |
-
# Optional detail layers, fetched by scripts/fetch_detail_geo.py. Without them
|
| 49 |
-
# the map shows outlines and major cities only, which is enough to place a
|
| 50 |
-
# country but not to aim within a city — while the human's map is street-level.
|
| 51 |
-
# The reward measures precision, so the two views have to agree about how
|
| 52 |
-
# precisely a pin can be aimed.
|
| 53 |
-
DETAIL_DIR = GEO_DIR / "detail"
|
| 54 |
-
|
| 55 |
-
# Detail appears as a function of zoom alone, never of proximity to the answer.
|
| 56 |
-
# Loading finer data near task locations would turn the cache into an oracle.
|
| 57 |
-
# Below this the map fetches real OSM ways, because Natural Earth 10m tops out
|
| 58 |
-
# at highway level: it will show the motorways around a city but not the street
|
| 59 |
-
# grid inside it, and the player's tiles show both. Responses are cached to
|
| 60 |
-
# disk, so the first render of a neighbourhood is slow and every later one is
|
| 61 |
-
# instant and byte-identical.
|
| 62 |
-
STREET_MAX_SPAN_DEG = 0.35
|
| 63 |
-
STREET_LABEL_MAX_SPAN_DEG = 0.08
|
| 64 |
-
"""Street names only below this span. Wider, and the names collide."""
|
| 65 |
-
MAX_STREET_LABELS = 14
|
| 66 |
-
"""A cap, because a dense grid has hundreds of named ways."""
|
| 67 |
-
OVERPASS_URL = "https://overpass-api.de/api/interpreter"
|
| 68 |
-
OVERPASS_TIMEOUT_S = 40.0
|
| 69 |
-
OSM_CACHE_DIR = pathlib.Path(
|
| 70 |
-
os.environ.get("GEOGUESSER_OSM_CACHE", str(GEO_DIR / "osm_cache"))
|
| 71 |
-
)
|
| 72 |
-
"""Where fetched street windows are cached.
|
| 73 |
-
|
| 74 |
-
Configurable so a deployment that cannot reach Overpass -- a Space, whose
|
| 75 |
-
datacenter egress gets 504s from the public instance -- can mount a pre-warmed
|
| 76 |
-
cache instead of silently rendering maps without streets.
|
| 77 |
-
"""
|
| 78 |
-
|
| 79 |
-
OVERPASS_ATTEMPTS = 2
|
| 80 |
-
"""Overpass 504s under load; one cheap retry recovers a useful fraction."""
|
| 81 |
-
|
| 82 |
-
_STREET_FETCH_FAILED = [False]
|
| 83 |
-
"""Whether the most recent street fetch failed, for observation metadata.
|
| 84 |
-
|
| 85 |
-
A map quietly missing its streets looks like a styling choice rather than a
|
| 86 |
-
degraded environment, so the failure is reported rather than inferred.
|
| 87 |
-
"""
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
def street_fetch_failed() -> bool:
|
| 91 |
-
"""Whether the last street fetch in this process failed."""
|
| 92 |
-
return _STREET_FETCH_FAILED[0]
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
_OSM_SCHEMA = 2
|
| 96 |
-
"""Bumped whenever the per-way record changes, to retire stale caches."""
|
| 97 |
-
|
| 98 |
-
# Ways worth drawing, heaviest first, with the line width to draw them at.
|
| 99 |
-
STREET_WEIGHTS = {
|
| 100 |
-
"motorway": 2.2,
|
| 101 |
-
"trunk": 2.0,
|
| 102 |
-
"primary": 1.7,
|
| 103 |
-
"secondary": 1.4,
|
| 104 |
-
"tertiary": 1.1,
|
| 105 |
-
"residential": 0.8,
|
| 106 |
-
"unclassified": 0.8,
|
| 107 |
-
"living_street": 0.7,
|
| 108 |
-
"service": 0.5,
|
| 109 |
-
"pedestrian": 0.5,
|
| 110 |
-
"motorway_link": 1.2,
|
| 111 |
-
"trunk_link": 1.1,
|
| 112 |
-
"primary_link": 1.0,
|
| 113 |
-
"secondary_link": 0.9,
|
| 114 |
-
"tertiary_link": 0.8,
|
| 115 |
-
}
|
| 116 |
-
|
| 117 |
-
URBAN_MAX_SPAN_DEG = 4.0
|
| 118 |
-
ROADS_MAX_SPAN_DEG = 4.0
|
| 119 |
-
RIVERS_MAX_SPAN_DEG = 6.0
|
| 120 |
-
TOWNS_MAX_SPAN_DEG = 4.0
|
| 121 |
-
|
| 122 |
-
# Version of the bundled geodata. Any change alters reverse-geocode text and
|
| 123 |
-
# therefore the observations, so eval scores must cite it.
|
| 124 |
-
GEODATA_VERSION = "natural-earth-110m+50m/2024.1"
|
| 125 |
-
|
| 126 |
-
_LAND = "#e9e5dd"
|
| 127 |
-
_ROAD = "#8a6a52"
|
| 128 |
-
_RIVER = "#6f9fba"
|
| 129 |
-
_URBAN = "#e4ded6"
|
| 130 |
-
_TOWN = "#3d3833"
|
| 131 |
-
_WATER = "#cfe0ea"
|
| 132 |
-
_BORDER = "#5f5a54"
|
| 133 |
-
_PIN = "#c4332a"
|
| 134 |
-
_TRUTH = "#3f7a55"
|
| 135 |
-
_INK = "#2b2a28"
|
| 136 |
-
|
| 137 |
-
_COMPASS = ("N", "NE", "E", "SE", "S", "SW", "W", "NW")
|
| 138 |
-
|
| 139 |
-
# Mutable so the server can turn street fetching off for a run that must never
|
| 140 |
-
# touch the network, without threading a flag through every render call.
|
| 141 |
-
_STREETS_ENABLED = [True]
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
def set_street_detail(enabled: bool) -> None:
|
| 145 |
-
"""Turn OSM street fetching on or off for this process."""
|
| 146 |
-
_STREETS_ENABLED[0] = bool(enabled)
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
def street_detail_enabled() -> bool:
|
| 150 |
-
"""Whether OSM street fetching is currently on."""
|
| 151 |
-
return _STREETS_ENABLED[0]
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
@dataclass(frozen=True)
|
| 155 |
-
class Place:
|
| 156 |
-
"""Where a coordinate falls, in words.
|
| 157 |
-
|
| 158 |
-
Attributes:
|
| 159 |
-
country (`str` or `None`):
|
| 160 |
-
Country name, or `None` over open water.
|
| 161 |
-
continent (`str` or `None`):
|
| 162 |
-
Continent name, or `None` over open water.
|
| 163 |
-
subregion (`str` or `None`):
|
| 164 |
-
UN subregion, or `None` over open water.
|
| 165 |
-
nearest_city (`str`):
|
| 166 |
-
Name of the closest populated place in the bundled dataset.
|
| 167 |
-
city_distance_km (`float`):
|
| 168 |
-
Distance to that city in kilometres.
|
| 169 |
-
city_bearing (`str`):
|
| 170 |
-
Compass direction from the city to the coordinate.
|
| 171 |
-
"""
|
| 172 |
-
|
| 173 |
-
country: str | None
|
| 174 |
-
continent: str | None
|
| 175 |
-
subregion: str | None
|
| 176 |
-
nearest_city: str
|
| 177 |
-
city_distance_km: float
|
| 178 |
-
city_bearing: str
|
| 179 |
-
|
| 180 |
-
|
| 181 |
-
@functools.lru_cache(maxsize=1)
|
| 182 |
-
def _countries() -> list[dict]:
|
| 183 |
-
return json.loads(COUNTRIES_FILE.read_text())["features"]
|
| 184 |
-
|
| 185 |
-
|
| 186 |
-
@functools.lru_cache(maxsize=1)
|
| 187 |
-
def _cities() -> list[tuple[str, float, float]]:
|
| 188 |
-
feats = json.loads(CITIES_FILE.read_text())["features"]
|
| 189 |
-
out = []
|
| 190 |
-
for f in feats:
|
| 191 |
-
lon, lat = f["geometry"]["coordinates"]
|
| 192 |
-
out.append((f["properties"]["name"], lat, lon))
|
| 193 |
-
return out
|
| 194 |
-
|
| 195 |
-
|
| 196 |
-
@functools.lru_cache(maxsize=8)
|
| 197 |
-
def _detail(layer: str) -> list[dict]:
|
| 198 |
-
"""Load one optional detail layer, or an empty list when absent."""
|
| 199 |
-
path = DETAIL_DIR / f"{layer}.json"
|
| 200 |
-
if not path.exists():
|
| 201 |
-
return []
|
| 202 |
-
return json.loads(path.read_text())
|
| 203 |
-
|
| 204 |
-
|
| 205 |
-
def _osm_cache_path(lat: float, lon: float, span: float) -> pathlib.Path:
|
| 206 |
-
"""Cache file for one street window, keyed by a quantised bounding box.
|
| 207 |
-
|
| 208 |
-
The key carries `_OSM_SCHEMA`, so widening what is stored per way retires
|
| 209 |
-
the old files instead of serving geometry with no labels.
|
| 210 |
-
"""
|
| 211 |
-
key = f"v{_OSM_SCHEMA}_{round(lat, 3):.3f}_{round(lon, 3):.3f}_{round(span, 4):.4f}"
|
| 212 |
-
digest = hashlib.sha1(key.encode()).hexdigest()[:16]
|
| 213 |
-
return OSM_CACHE_DIR / f"{digest}.json"
|
| 214 |
-
|
| 215 |
-
|
| 216 |
-
def _fetch_streets(lat: float, lon: float, span: float) -> list[dict]:
|
| 217 |
-
"""Ask Overpass for the ways in one window, or return an empty list.
|
| 218 |
-
|
| 219 |
-
Any failure — offline, rate limited, malformed — degrades to no streets
|
| 220 |
-
rather than failing the step that asked for the map.
|
| 221 |
-
"""
|
| 222 |
-
box = f"{lat - span},{lon - span},{lat + span},{lon + span}"
|
| 223 |
-
# Street names and road numbers are what make the agent's map comparable to
|
| 224 |
-
# the human's, and they cost nothing extra: Overpass returns tags with the
|
| 225 |
-
# geometry either way.
|
| 226 |
-
query = f'[out:json][timeout:30];way["highway"]({box});out tags geom;'
|
| 227 |
-
payload = None
|
| 228 |
-
for attempt in range(1, OVERPASS_ATTEMPTS + 1):
|
| 229 |
-
request = urllib.request.Request(
|
| 230 |
-
OVERPASS_URL,
|
| 231 |
-
data=urllib.parse.urlencode({"data": query}).encode(),
|
| 232 |
-
headers={"User-Agent": "openenv-geoguesser-env/0.1 (research environment)"},
|
| 233 |
-
)
|
| 234 |
-
try:
|
| 235 |
-
with urllib.request.urlopen(
|
| 236 |
-
request, timeout=OVERPASS_TIMEOUT_S
|
| 237 |
-
) as response:
|
| 238 |
-
payload = json.loads(response.read())
|
| 239 |
-
break
|
| 240 |
-
except Exception as exc: # noqa: BLE001 - a missing map must not end a step
|
| 241 |
-
logger.warning(
|
| 242 |
-
"street detail unavailable for %.3f,%.3f (attempt %d/%d): %r",
|
| 243 |
-
lat,
|
| 244 |
-
lon,
|
| 245 |
-
attempt,
|
| 246 |
-
OVERPASS_ATTEMPTS,
|
| 247 |
-
exc,
|
| 248 |
-
)
|
| 249 |
-
if payload is None:
|
| 250 |
-
_STREET_FETCH_FAILED[0] = True
|
| 251 |
-
return []
|
| 252 |
-
_STREET_FETCH_FAILED[0] = False
|
| 253 |
-
ways = []
|
| 254 |
-
for element in payload.get("elements", []):
|
| 255 |
-
geometry = element.get("geometry") or []
|
| 256 |
-
if len(geometry) < 2:
|
| 257 |
-
continue
|
| 258 |
-
tags = element.get("tags") or {}
|
| 259 |
-
kind = tags.get("highway", "residential")
|
| 260 |
-
way = {
|
| 261 |
-
"w": STREET_WEIGHTS.get(kind, 0.6),
|
| 262 |
-
"c": [[point["lon"], point["lat"]] for point in geometry],
|
| 263 |
-
}
|
| 264 |
-
# A road number is often the only label a rural road has, and it is
|
| 265 |
-
# exactly the clue a player reads off a sign.
|
| 266 |
-
label = tags.get("name") or tags.get("ref")
|
| 267 |
-
if label:
|
| 268 |
-
way["n"] = label[:34]
|
| 269 |
-
ways.append(way)
|
| 270 |
-
return ways
|
| 271 |
-
|
| 272 |
-
|
| 273 |
-
def street_ways(lat: float, lon: float, span: float) -> list[dict]:
|
| 274 |
-
"""
|
| 275 |
-
Street geometry for one window, cached on disk.
|
| 276 |
-
|
| 277 |
-
Args:
|
| 278 |
-
lat (`float`):
|
| 279 |
-
Latitude at the centre of the window.
|
| 280 |
-
lon (`float`):
|
| 281 |
-
Longitude at the centre of the window.
|
| 282 |
-
span (`float`):
|
| 283 |
-
Half-width of the window in degrees.
|
| 284 |
-
|
| 285 |
-
Returns:
|
| 286 |
-
`list[dict]`: One entry per way, with a line width `w` and a coordinate
|
| 287 |
-
list `c`. Empty when the window is too wide, streets are disabled, or
|
| 288 |
-
the fetch failed.
|
| 289 |
-
"""
|
| 290 |
-
path = _osm_cache_path(lat, lon, span)
|
| 291 |
-
if path.exists():
|
| 292 |
-
try:
|
| 293 |
-
return json.loads(path.read_text())
|
| 294 |
-
except json.JSONDecodeError:
|
| 295 |
-
path.unlink(missing_ok=True)
|
| 296 |
-
ways = _fetch_streets(lat, lon, span)
|
| 297 |
-
OSM_CACHE_DIR.mkdir(parents=True, exist_ok=True)
|
| 298 |
-
path.write_text(json.dumps(ways, separators=(",", ":")))
|
| 299 |
-
logger.info(
|
| 300 |
-
"cached %d street ways for %.3f,%.3f span %.3f", len(ways), lat, lon, span
|
| 301 |
-
)
|
| 302 |
-
return ways
|
| 303 |
-
|
| 304 |
-
|
| 305 |
-
def has_detail() -> bool:
|
| 306 |
-
"""Whether the optional detail layers are installed."""
|
| 307 |
-
return any(
|
| 308 |
-
(DETAIL_DIR / f"{name}.json").exists()
|
| 309 |
-
for name in ("roads", "urban", "places", "rivers")
|
| 310 |
-
)
|
| 311 |
-
|
| 312 |
-
|
| 313 |
-
def _parts(row: dict) -> list[list]:
|
| 314 |
-
"""Coordinate lists for one compacted feature, whatever its geometry."""
|
| 315 |
-
kind, coordinates = row["t"], row["c"]
|
| 316 |
-
if kind in ("LineString", "Point"):
|
| 317 |
-
return [coordinates] if kind == "LineString" else [[coordinates]]
|
| 318 |
-
if kind == "MultiLineString":
|
| 319 |
-
return coordinates
|
| 320 |
-
if kind == "Polygon":
|
| 321 |
-
return [coordinates[0]]
|
| 322 |
-
if kind == "MultiPolygon":
|
| 323 |
-
return [polygon[0] for polygon in coordinates]
|
| 324 |
-
return []
|
| 325 |
-
|
| 326 |
-
|
| 327 |
-
def _near(points: list, lat: float, lon: float, span: float) -> bool:
|
| 328 |
-
return any(
|
| 329 |
-
abs(point[0] - lon) < span and abs(point[1] - lat) < span for point in points
|
| 330 |
-
)
|
| 331 |
-
|
| 332 |
-
|
| 333 |
-
def _rings(geometry: dict) -> list[list]:
|
| 334 |
-
if geometry["type"] == "Polygon":
|
| 335 |
-
return [geometry["coordinates"][0]]
|
| 336 |
-
if geometry["type"] == "MultiPolygon":
|
| 337 |
-
return [poly[0] for poly in geometry["coordinates"]]
|
| 338 |
-
return []
|
| 339 |
-
|
| 340 |
-
|
| 341 |
-
def _bearing(from_lat: float, from_lon: float, to_lat: float, to_lon: float) -> str:
|
| 342 |
-
d_lon = math.radians(to_lon - from_lon)
|
| 343 |
-
lat_a, lat_b = math.radians(from_lat), math.radians(to_lat)
|
| 344 |
-
y = math.sin(d_lon) * math.cos(lat_b)
|
| 345 |
-
x = math.cos(lat_a) * math.sin(lat_b) - math.sin(lat_a) * math.cos(
|
| 346 |
-
lat_b
|
| 347 |
-
) * math.cos(d_lon)
|
| 348 |
-
deg = (math.degrees(math.atan2(y, x)) + 360) % 360
|
| 349 |
-
return _COMPASS[int(deg / 45 + 0.5) % 8]
|
| 350 |
-
|
| 351 |
-
|
| 352 |
-
def locate(lat: float, lon: float) -> Place:
|
| 353 |
-
"""
|
| 354 |
-
Describe a coordinate using only the bundled vectors.
|
| 355 |
-
|
| 356 |
-
Args:
|
| 357 |
-
lat (`float`):
|
| 358 |
-
Latitude in degrees.
|
| 359 |
-
lon (`float`):
|
| 360 |
-
Longitude in degrees.
|
| 361 |
-
|
| 362 |
-
Returns:
|
| 363 |
-
[`Place`]: Administrative names for the point, plus the nearest
|
| 364 |
-
populated place with distance and bearing.
|
| 365 |
-
"""
|
| 366 |
-
country = continent = subregion = None
|
| 367 |
-
for feature in _countries():
|
| 368 |
-
for ring in _rings(feature["geometry"]):
|
| 369 |
-
if MplPath(ring).contains_point((lon, lat)):
|
| 370 |
-
props = feature["properties"]
|
| 371 |
-
country = props.get("ADMIN")
|
| 372 |
-
continent = props.get("CONTINENT")
|
| 373 |
-
subregion = props.get("SUBREGION")
|
| 374 |
-
break
|
| 375 |
-
if country:
|
| 376 |
-
break
|
| 377 |
-
|
| 378 |
-
best_name, best_km, best_bearing = "", float("inf"), "N"
|
| 379 |
-
for name, city_lat, city_lon in _cities():
|
| 380 |
-
km = _haversine_km(lat, lon, city_lat, city_lon)
|
| 381 |
-
if km < best_km:
|
| 382 |
-
best_name, best_km = name, km
|
| 383 |
-
best_bearing = _bearing(city_lat, city_lon, lat, lon)
|
| 384 |
-
return Place(country, continent, subregion, best_name, best_km, best_bearing)
|
| 385 |
-
|
| 386 |
-
|
| 387 |
-
def describe_pin(
|
| 388 |
-
index: int, lat: float, lon: float, previous: tuple[float, float] | None = None
|
| 389 |
-
) -> str:
|
| 390 |
-
"""
|
| 391 |
-
One line of feedback for a pin, containing nothing about the target.
|
| 392 |
-
|
| 393 |
-
Args:
|
| 394 |
-
index (`int`):
|
| 395 |
-
1-based pin number.
|
| 396 |
-
lat (`float`):
|
| 397 |
-
Latitude of the pin.
|
| 398 |
-
lon (`float`):
|
| 399 |
-
Longitude of the pin.
|
| 400 |
-
previous (`tuple[float, float]`, *optional*):
|
| 401 |
-
The preceding pin, so the agent can compare its own candidates.
|
| 402 |
-
|
| 403 |
-
Returns:
|
| 404 |
-
`str`: Feedback describing the pinned location.
|
| 405 |
-
"""
|
| 406 |
-
place = locate(lat, lon)
|
| 407 |
-
where = (
|
| 408 |
-
f"{place.country} ({place.subregion})"
|
| 409 |
-
if place.country
|
| 410 |
-
else "open water - no landmass at this coordinate"
|
| 411 |
-
)
|
| 412 |
-
line = (
|
| 413 |
-
f"Pin {index} placed at {lat:.4f}, {lon:.4f} - {where}. "
|
| 414 |
-
f"Nearest major city: {place.nearest_city}, "
|
| 415 |
-
f"~{place.city_distance_km:.0f} km {place.city_bearing}."
|
| 416 |
-
)
|
| 417 |
-
if previous is not None:
|
| 418 |
-
km = _haversine_km(lat, lon, previous[0], previous[1])
|
| 419 |
-
line += f" Distance from pin {index - 1}: {km:.0f} km."
|
| 420 |
-
return line
|
| 421 |
-
|
| 422 |
-
|
| 423 |
-
def _draw_land(ax, linewidth: float) -> None:
|
| 424 |
-
for feature in _countries():
|
| 425 |
-
for ring in _rings(feature["geometry"]):
|
| 426 |
-
ax.add_patch(
|
| 427 |
-
MplPolygon(
|
| 428 |
-
ring,
|
| 429 |
-
closed=True,
|
| 430 |
-
facecolor=_LAND,
|
| 431 |
-
edgecolor=_BORDER,
|
| 432 |
-
linewidth=linewidth,
|
| 433 |
-
)
|
| 434 |
-
)
|
| 435 |
-
|
| 436 |
-
|
| 437 |
-
def _label_countries(ax, lat: float, lon: float, span: float) -> None:
|
| 438 |
-
"""Label countries visible in the window, clipped to the viewport."""
|
| 439 |
-
for feature in _countries():
|
| 440 |
-
for ring in _rings(feature["geometry"]):
|
| 441 |
-
visible = [
|
| 442 |
-
p for p in ring if abs(p[0] - lon) < span and abs(p[1] - lat) < span
|
| 443 |
-
]
|
| 444 |
-
if len(visible) > 3:
|
| 445 |
-
cx = sum(p[0] for p in visible) / len(visible)
|
| 446 |
-
cy = sum(p[1] for p in visible) / len(visible)
|
| 447 |
-
ax.text(
|
| 448 |
-
cx,
|
| 449 |
-
cy,
|
| 450 |
-
feature["properties"]["NAME"],
|
| 451 |
-
fontsize=7.4,
|
| 452 |
-
ha="center",
|
| 453 |
-
color="#55504a",
|
| 454 |
-
zorder=7,
|
| 455 |
-
)
|
| 456 |
-
break
|
| 457 |
-
|
| 458 |
-
|
| 459 |
-
def _draw_detail(ax, lat: float, lon: float, span: float) -> None:
|
| 460 |
-
"""Draw urban areas, rivers, roads and town names, by zoom level.
|
| 461 |
-
|
| 462 |
-
Each layer switches on below its own span so a wide view stays legible and a
|
| 463 |
-
tight view shows enough road structure to aim a pin within a town.
|
| 464 |
-
"""
|
| 465 |
-
if span <= URBAN_MAX_SPAN_DEG:
|
| 466 |
-
for row in _detail("urban"):
|
| 467 |
-
for part in _parts(row):
|
| 468 |
-
if _near(part, lat, lon, span):
|
| 469 |
-
ax.add_patch(
|
| 470 |
-
MplPolygon(
|
| 471 |
-
part,
|
| 472 |
-
closed=True,
|
| 473 |
-
facecolor=_URBAN,
|
| 474 |
-
edgecolor="none",
|
| 475 |
-
zorder=2,
|
| 476 |
-
)
|
| 477 |
-
)
|
| 478 |
-
if span <= RIVERS_MAX_SPAN_DEG:
|
| 479 |
-
segments = [
|
| 480 |
-
part
|
| 481 |
-
for row in _detail("rivers")
|
| 482 |
-
for part in _parts(row)
|
| 483 |
-
if _near(part, lat, lon, span)
|
| 484 |
-
]
|
| 485 |
-
if segments:
|
| 486 |
-
ax.add_collection(
|
| 487 |
-
LineCollection(segments, colors=_RIVER, linewidths=0.9, zorder=3)
|
| 488 |
-
)
|
| 489 |
-
if span <= ROADS_MAX_SPAN_DEG:
|
| 490 |
-
segments = [
|
| 491 |
-
part
|
| 492 |
-
for row in _detail("roads")
|
| 493 |
-
for part in _parts(row)
|
| 494 |
-
if _near(part, lat, lon, span)
|
| 495 |
-
]
|
| 496 |
-
if segments:
|
| 497 |
-
ax.add_collection(
|
| 498 |
-
LineCollection(segments, colors=_ROAD, linewidths=1.4, zorder=4)
|
| 499 |
-
)
|
| 500 |
-
if span <= STREET_MAX_SPAN_DEG and _STREETS_ENABLED[0]:
|
| 501 |
-
ways = street_ways(lat, lon, span)
|
| 502 |
-
# Generalise by zoom the way a real style does: drawing every service
|
| 503 |
-
# road and footpath at city scale turns the grid into a smear, so the
|
| 504 |
-
# minor classes only appear once the window is tight enough to hold
|
| 505 |
-
# them, and widths grow as the window shrinks.
|
| 506 |
-
floor = 0.75 if span > 0.15 else (0.55 if span > 0.05 else 0.0)
|
| 507 |
-
scale = 1.0 if span > 0.15 else (1.4 if span > 0.05 else 1.9)
|
| 508 |
-
drawn = [way for way in ways if way["w"] >= floor]
|
| 509 |
-
for weight in sorted({way["w"] for way in drawn}):
|
| 510 |
-
segments = [way["c"] for way in drawn if way["w"] == weight]
|
| 511 |
-
width = weight * scale
|
| 512 |
-
# A casing under a white fill is what makes a dense grid legible; a
|
| 513 |
-
# single flat colour reads as noise.
|
| 514 |
-
ax.add_collection(
|
| 515 |
-
LineCollection(
|
| 516 |
-
segments, colors="#c9bfb3", linewidths=width + 0.5, zorder=4
|
| 517 |
-
)
|
| 518 |
-
)
|
| 519 |
-
ax.add_collection(
|
| 520 |
-
LineCollection(segments, colors="#ffffff", linewidths=width, zorder=5)
|
| 521 |
-
)
|
| 522 |
-
_label_streets(ax, drawn, lat, lon, span)
|
| 523 |
-
if span <= TOWNS_MAX_SPAN_DEG:
|
| 524 |
-
shown = 0
|
| 525 |
-
for row in _detail("places"):
|
| 526 |
-
point = row["c"]
|
| 527 |
-
if abs(point[0] - lon) < span * 0.95 and abs(point[1] - lat) < span * 0.95:
|
| 528 |
-
ax.plot(
|
| 529 |
-
point[0],
|
| 530 |
-
point[1],
|
| 531 |
-
"o",
|
| 532 |
-
markersize=2.6,
|
| 533 |
-
markerfacecolor=_TOWN,
|
| 534 |
-
markeredgecolor="none",
|
| 535 |
-
zorder=6,
|
| 536 |
-
)
|
| 537 |
-
if row.get("n"):
|
| 538 |
-
label = ax.text(
|
| 539 |
-
point[0],
|
| 540 |
-
point[1] + span * 0.03,
|
| 541 |
-
row["n"],
|
| 542 |
-
fontsize=6.2,
|
| 543 |
-
ha="center",
|
| 544 |
-
color=_TOWN,
|
| 545 |
-
zorder=9,
|
| 546 |
-
)
|
| 547 |
-
label.set_path_effects(
|
| 548 |
-
[
|
| 549 |
-
path_effects.Stroke(linewidth=1.8, foreground="#ffffff"),
|
| 550 |
-
path_effects.Normal(),
|
| 551 |
-
]
|
| 552 |
-
)
|
| 553 |
-
shown += 1
|
| 554 |
-
if shown >= (8 if span <= STREET_MAX_SPAN_DEG else 28):
|
| 555 |
-
break
|
| 556 |
-
|
| 557 |
-
|
| 558 |
-
def _label_streets(ax, ways: list[dict], lat: float, lon: float, span: float) -> None:
|
| 559 |
-
"""
|
| 560 |
-
Write street names along the ways, the way a real map style does.
|
| 561 |
-
|
| 562 |
-
One label per name, on that name's longest visible run, rotated to follow
|
| 563 |
-
the road and haloed so it stays readable over the casing. Longest-run
|
| 564 |
-
selection matters: labelling an arbitrary segment puts "Main Street" on a
|
| 565 |
-
50 m stub while the avenue itself goes unnamed.
|
| 566 |
-
|
| 567 |
-
Args:
|
| 568 |
-
ax:
|
| 569 |
-
Matplotlib axes to draw on.
|
| 570 |
-
ways (`list[dict]`):
|
| 571 |
-
Way records from [`street_ways`], some carrying a name in `n`.
|
| 572 |
-
lat (`float`):
|
| 573 |
-
Latitude at the centre of the window.
|
| 574 |
-
lon (`float`):
|
| 575 |
-
Longitude at the centre of the window.
|
| 576 |
-
span (`float`):
|
| 577 |
-
Half-width of the window in degrees.
|
| 578 |
-
"""
|
| 579 |
-
if span > STREET_LABEL_MAX_SPAN_DEG:
|
| 580 |
-
return
|
| 581 |
-
# Pick, per name, the longest run of points that actually falls inside the
|
| 582 |
-
# window, so the label lands where the reader can see it.
|
| 583 |
-
best: dict[str, tuple[float, list]] = {}
|
| 584 |
-
for way in ways:
|
| 585 |
-
name = way.get("n")
|
| 586 |
-
if not name:
|
| 587 |
-
continue
|
| 588 |
-
inside = [
|
| 589 |
-
point
|
| 590 |
-
for point in way["c"]
|
| 591 |
-
if abs(point[0] - lon) < span * 0.92 and abs(point[1] - lat) < span * 0.92
|
| 592 |
-
]
|
| 593 |
-
if len(inside) < 2:
|
| 594 |
-
continue
|
| 595 |
-
length = sum(
|
| 596 |
-
math.dist(inside[i], inside[i + 1]) for i in range(len(inside) - 1)
|
| 597 |
-
)
|
| 598 |
-
if name not in best or length > best[name][0]:
|
| 599 |
-
best[name] = (length, inside)
|
| 600 |
-
|
| 601 |
-
ordered = sorted(best.items(), key=lambda item: -item[1][0])
|
| 602 |
-
aspect = max(0.05, math.cos(math.radians(lat)))
|
| 603 |
-
# Parallel streets in a grid all have their midpoint in the same place, so
|
| 604 |
-
# placing every label at its midpoint stacks them into an unreadable pile.
|
| 605 |
-
# Claim one coarse cell per label, walking along each road to find a free
|
| 606 |
-
# one, and drop the label rather than overprint. Longest roads go first, so
|
| 607 |
-
# the ones worth naming win the space.
|
| 608 |
-
occupied: set[tuple[int, int]] = set()
|
| 609 |
-
# The cell has to be taller than the crowding you want to break up, not
|
| 610 |
-
# just taller than the glyphs: parallel streets one block apart land in
|
| 611 |
-
# different fine cells and still read as a stack. A tall cell forces
|
| 612 |
-
# neighbours to slide along their own road instead, which is what a real
|
| 613 |
-
# map style does.
|
| 614 |
-
cell_x = span * 0.22
|
| 615 |
-
cell_y = span * 0.13
|
| 616 |
-
placed = 0
|
| 617 |
-
for name, (_, points) in ordered:
|
| 618 |
-
if placed >= MAX_STREET_LABELS:
|
| 619 |
-
break
|
| 620 |
-
middle = len(points) // 2
|
| 621 |
-
# Try the midpoint first, then positions either side of it.
|
| 622 |
-
order = sorted(range(len(points)), key=lambda i: abs(i - middle))
|
| 623 |
-
chosen = None
|
| 624 |
-
for i in order:
|
| 625 |
-
cell = (
|
| 626 |
-
int((points[i][0] - lon) / cell_x),
|
| 627 |
-
int((points[i][1] - lat) / cell_y),
|
| 628 |
-
)
|
| 629 |
-
if cell not in occupied:
|
| 630 |
-
occupied.add(cell)
|
| 631 |
-
chosen = i
|
| 632 |
-
break
|
| 633 |
-
if chosen is None:
|
| 634 |
-
continue
|
| 635 |
-
placed += 1
|
| 636 |
-
middle = chosen
|
| 637 |
-
start = points[max(0, middle - 1)]
|
| 638 |
-
end = points[min(len(points) - 1, middle + 1)]
|
| 639 |
-
# Longitude degrees are shorter than latitude ones away from the
|
| 640 |
-
# equator, so the on-screen angle needs the cos(lat) correction or
|
| 641 |
-
# labels sit visibly off their road.
|
| 642 |
-
angle = math.degrees(
|
| 643 |
-
math.atan2(end[1] - start[1], (end[0] - start[0]) * aspect)
|
| 644 |
-
)
|
| 645 |
-
if angle > 90:
|
| 646 |
-
angle -= 180
|
| 647 |
-
elif angle < -90:
|
| 648 |
-
angle += 180
|
| 649 |
-
text = ax.text(
|
| 650 |
-
points[middle][0],
|
| 651 |
-
points[middle][1],
|
| 652 |
-
name,
|
| 653 |
-
fontsize=4.6,
|
| 654 |
-
ha="center",
|
| 655 |
-
va="center",
|
| 656 |
-
rotation=angle,
|
| 657 |
-
rotation_mode="anchor",
|
| 658 |
-
color="#4a453f",
|
| 659 |
-
zorder=8,
|
| 660 |
-
)
|
| 661 |
-
text.set_path_effects(
|
| 662 |
-
[
|
| 663 |
-
path_effects.Stroke(linewidth=1.4, foreground="#ffffff"),
|
| 664 |
-
path_effects.Normal(),
|
| 665 |
-
]
|
| 666 |
-
)
|
| 667 |
-
|
| 668 |
-
|
| 669 |
-
def _scale_bar(ax, lat: float, lon: float, span: float) -> None:
|
| 670 |
-
km_per_degree = 111.32 * max(0.05, math.cos(math.radians(lat)))
|
| 671 |
-
# A bar reading "100 km" across a hemisphere is worse than no bar.
|
| 672 |
-
if span > 20:
|
| 673 |
-
unit = 2000
|
| 674 |
-
elif span > 5:
|
| 675 |
-
unit = 500
|
| 676 |
-
elif span > 2:
|
| 677 |
-
unit = 100
|
| 678 |
-
elif span > 0.5:
|
| 679 |
-
unit = 20
|
| 680 |
-
else:
|
| 681 |
-
unit = 5
|
| 682 |
-
bar = unit / km_per_degree
|
| 683 |
-
x0, y0 = lon - span * 0.9, lat - span * 0.9
|
| 684 |
-
ax.plot([x0, x0 + bar], [y0, y0], color=_INK, linewidth=2.2, zorder=9)
|
| 685 |
-
ax.text(
|
| 686 |
-
x0 + bar / 2,
|
| 687 |
-
y0 + span * 0.04,
|
| 688 |
-
f"{unit} km",
|
| 689 |
-
fontsize=6.6,
|
| 690 |
-
ha="center",
|
| 691 |
-
color=_INK,
|
| 692 |
-
zorder=9,
|
| 693 |
-
)
|
| 694 |
-
|
| 695 |
-
|
| 696 |
-
def _plot_pins(ax, pins: list[tuple[float, float]], size: float) -> None:
|
| 697 |
-
for i, (lat, lon) in enumerate(pins, 1):
|
| 698 |
-
ax.plot(
|
| 699 |
-
lon,
|
| 700 |
-
lat,
|
| 701 |
-
marker="o",
|
| 702 |
-
markersize=size,
|
| 703 |
-
markerfacecolor=_PIN,
|
| 704 |
-
markeredgecolor="white",
|
| 705 |
-
markeredgewidth=1.5,
|
| 706 |
-
zorder=8,
|
| 707 |
-
)
|
| 708 |
-
ax.annotate(
|
| 709 |
-
str(i),
|
| 710 |
-
(lon, lat),
|
| 711 |
-
color="white",
|
| 712 |
-
fontsize=size * 0.62,
|
| 713 |
-
weight="bold",
|
| 714 |
-
ha="center",
|
| 715 |
-
va="center",
|
| 716 |
-
zorder=9,
|
| 717 |
-
)
|
| 718 |
-
|
| 719 |
-
|
| 720 |
-
def _plot_truth(ax, pins, truth: tuple[float, float], size: float) -> None:
|
| 721 |
-
"""Draw the true location and the line to the guess it is being compared to."""
|
| 722 |
-
truth_lat, truth_lon = truth
|
| 723 |
-
if pins:
|
| 724 |
-
guess_lat, guess_lon = pins[-1]
|
| 725 |
-
ax.plot(
|
| 726 |
-
[guess_lon, truth_lon],
|
| 727 |
-
[guess_lat, truth_lat],
|
| 728 |
-
linestyle="--",
|
| 729 |
-
color=_PIN,
|
| 730 |
-
linewidth=1.6,
|
| 731 |
-
zorder=7,
|
| 732 |
-
)
|
| 733 |
-
ax.plot(
|
| 734 |
-
truth_lon,
|
| 735 |
-
truth_lat,
|
| 736 |
-
marker="o",
|
| 737 |
-
markersize=size,
|
| 738 |
-
markerfacecolor=_TRUTH,
|
| 739 |
-
markeredgecolor="white",
|
| 740 |
-
markeredgewidth=1.5,
|
| 741 |
-
zorder=10,
|
| 742 |
-
)
|
| 743 |
-
|
| 744 |
-
|
| 745 |
-
def render_map(
|
| 746 |
-
pins: list[tuple[float, float]],
|
| 747 |
-
focus: tuple[float, float] | None = None,
|
| 748 |
-
span_deg: float = 7.0,
|
| 749 |
-
dpi: int = 100,
|
| 750 |
-
truth: tuple[float, float] | None = None,
|
| 751 |
-
) -> Image.Image:
|
| 752 |
-
"""
|
| 753 |
-
Render the guess map: a world panel plus a zoomed panel.
|
| 754 |
-
|
| 755 |
-
Args:
|
| 756 |
-
pins (`list[tuple[float, float]]`):
|
| 757 |
-
Pins as `(lat, lon)`, drawn and numbered in order.
|
| 758 |
-
focus (`tuple[float, float]`, *optional*):
|
| 759 |
-
Centre of the zoomed panel. Defaults to the last pin.
|
| 760 |
-
span_deg (`float`, *optional*, defaults to `7.0`):
|
| 761 |
-
Half-width of the zoomed panel in degrees. Below roughly 4 degrees
|
| 762 |
-
the panel adds urban areas, roads, rivers and town names, when the
|
| 763 |
-
optional detail layers are installed.
|
| 764 |
-
dpi (`int`, *optional*, defaults to `100`):
|
| 765 |
-
Figure resolution.
|
| 766 |
-
truth (`tuple[float, float]`, *optional*):
|
| 767 |
-
Ground truth, drawn in green with a dashed line to the last pin.
|
| 768 |
-
Only ever passed after a guess has been scored, so it cannot leak
|
| 769 |
-
into an observation the agent sees before committing.
|
| 770 |
-
|
| 771 |
-
Returns:
|
| 772 |
-
`PIL.Image.Image`: The two-panel map.
|
| 773 |
-
"""
|
| 774 |
-
fig, (world, zoom) = plt.subplots(
|
| 775 |
-
1, 2, figsize=(10.2, 3.9), dpi=dpi, gridspec_kw={"width_ratios": [1.55, 1]}
|
| 776 |
-
)
|
| 777 |
-
|
| 778 |
-
_draw_land(world, 0.35)
|
| 779 |
-
world.set_xlim(-180, 180)
|
| 780 |
-
world.set_ylim(-90, 90)
|
| 781 |
-
world.set_facecolor(_WATER)
|
| 782 |
-
for x in range(-180, 181, 60):
|
| 783 |
-
world.axvline(x, color="white", linewidth=0.5, alpha=0.7)
|
| 784 |
-
for y in range(-60, 61, 30):
|
| 785 |
-
world.axhline(y, color="white", linewidth=0.5, alpha=0.7)
|
| 786 |
-
_plot_pins(world, pins, 8.0)
|
| 787 |
-
if truth is not None:
|
| 788 |
-
_plot_truth(world, pins, truth, 8.0)
|
| 789 |
-
title = "guess and true location" if truth is not None else "your pins - world"
|
| 790 |
-
world.set_title(title, fontsize=9, loc="left", color="#333")
|
| 791 |
-
world.set_xticks([])
|
| 792 |
-
world.set_yticks([])
|
| 793 |
-
|
| 794 |
-
focus_lat, focus_lon = focus if focus else (pins[-1] if pins else (20.0, 0.0))
|
| 795 |
-
_draw_land(zoom, 0.9)
|
| 796 |
-
zoom.set_xlim(focus_lon - span_deg, focus_lon + span_deg)
|
| 797 |
-
zoom.set_ylim(focus_lat - span_deg, focus_lat + span_deg)
|
| 798 |
-
zoom.set_facecolor(_WATER)
|
| 799 |
-
_draw_detail(zoom, focus_lat, focus_lon, span_deg)
|
| 800 |
-
# Above about 20 degrees every country in a hemisphere wants a label and the
|
| 801 |
-
# panel turns into a stack of overlapping text.
|
| 802 |
-
if TOWNS_MAX_SPAN_DEG < span_deg <= 20.0:
|
| 803 |
-
_label_countries(zoom, focus_lat, focus_lon, span_deg)
|
| 804 |
-
_plot_pins(zoom, pins, 10.0)
|
| 805 |
-
if truth is not None:
|
| 806 |
-
_plot_truth(zoom, pins, truth, 10.0)
|
| 807 |
-
_scale_bar(zoom, focus_lat, focus_lon, span_deg)
|
| 808 |
-
# Below a degree, degrees round to "0"; kilometres are the useful unit there.
|
| 809 |
-
width_deg = span_deg * 2
|
| 810 |
-
label = (
|
| 811 |
-
f"{width_deg:.0f} deg view"
|
| 812 |
-
if width_deg >= 1
|
| 813 |
-
else f"{width_deg * 111:.0f} km view"
|
| 814 |
-
)
|
| 815 |
-
zoom.set_title(label, fontsize=9, loc="left", color="#333")
|
| 816 |
-
zoom.set_xticks([])
|
| 817 |
-
zoom.set_yticks([])
|
| 818 |
-
|
| 819 |
-
fig.tight_layout(pad=0.6)
|
| 820 |
-
buf = io.BytesIO()
|
| 821 |
-
fig.savefig(buf, format="png", facecolor="white")
|
| 822 |
-
plt.close(fig)
|
| 823 |
-
buf.seek(0)
|
| 824 |
-
return Image.open(buf).convert("RGB")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/render/pano.py
DELETED
|
@@ -1,108 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Perspective views out of an equirectangular panorama.
|
| 4 |
-
|
| 5 |
-
`look()` is a gnomonic reprojection: build a camera ray for every output
|
| 6 |
-
pixel, rotate it by the requested heading and pitch, convert to spherical
|
| 7 |
-
coordinates and sample the source image. Integer sampling keeps it
|
| 8 |
-
deterministic — the same arguments always produce the same bytes, which is
|
| 9 |
-
what lets a GRPO group share one starting observation.
|
| 10 |
-
"""
|
| 11 |
-
|
| 12 |
-
from __future__ import annotations
|
| 13 |
-
|
| 14 |
-
import base64
|
| 15 |
-
import io
|
| 16 |
-
|
| 17 |
-
import numpy as np
|
| 18 |
-
from PIL import Image
|
| 19 |
-
|
| 20 |
-
|
| 21 |
-
DEFAULT_SIZE = (640, 640)
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
def look(
|
| 25 |
-
pano: Image.Image,
|
| 26 |
-
heading_deg: float,
|
| 27 |
-
pitch_deg: float = 0.0,
|
| 28 |
-
fov_deg: float = 90.0,
|
| 29 |
-
size: tuple[int, int] = DEFAULT_SIZE,
|
| 30 |
-
) -> Image.Image:
|
| 31 |
-
"""
|
| 32 |
-
Render one perspective view out of an equirectangular panorama.
|
| 33 |
-
|
| 34 |
-
Args:
|
| 35 |
-
pano (`PIL.Image.Image`):
|
| 36 |
-
Source panorama, 2:1 equirectangular.
|
| 37 |
-
heading_deg (`float`):
|
| 38 |
-
Compass heading in degrees, `0` being the panorama's own north.
|
| 39 |
-
pitch_deg (`float`, *optional*, defaults to `0.0`):
|
| 40 |
-
Vertical angle in degrees; positive looks up.
|
| 41 |
-
fov_deg (`float`, *optional*, defaults to `90.0`):
|
| 42 |
-
Horizontal field of view. Smaller values zoom in.
|
| 43 |
-
size (`tuple[int, int]`, *optional*, defaults to `(640, 640)`):
|
| 44 |
-
Output width and height in pixels.
|
| 45 |
-
|
| 46 |
-
Returns:
|
| 47 |
-
`PIL.Image.Image`: The rendered view.
|
| 48 |
-
|
| 49 |
-
Examples:
|
| 50 |
-
|
| 51 |
-
```python
|
| 52 |
-
view = look(Image.open("pano.jpg"), heading_deg=90, fov_deg=30)
|
| 53 |
-
```
|
| 54 |
-
"""
|
| 55 |
-
src = np.asarray(pano.convert("RGB"))
|
| 56 |
-
src_h, src_w = src.shape[:2]
|
| 57 |
-
width, height = size
|
| 58 |
-
|
| 59 |
-
focal = 0.5 * width / np.tan(np.radians(fov_deg) / 2)
|
| 60 |
-
xs, ys = np.meshgrid(np.arange(width) - width / 2, np.arange(height) - height / 2)
|
| 61 |
-
rays = np.stack([xs, -ys, np.full_like(xs, focal, dtype=float)], axis=-1)
|
| 62 |
-
rays /= np.linalg.norm(rays, axis=-1, keepdims=True)
|
| 63 |
-
|
| 64 |
-
pitch, heading = np.radians(pitch_deg), np.radians(heading_deg)
|
| 65 |
-
rot_x = np.array(
|
| 66 |
-
[
|
| 67 |
-
[1, 0, 0],
|
| 68 |
-
[0, np.cos(pitch), -np.sin(pitch)],
|
| 69 |
-
[0, np.sin(pitch), np.cos(pitch)],
|
| 70 |
-
]
|
| 71 |
-
)
|
| 72 |
-
rot_y = np.array(
|
| 73 |
-
[
|
| 74 |
-
[np.cos(heading), 0, np.sin(heading)],
|
| 75 |
-
[0, 1, 0],
|
| 76 |
-
[-np.sin(heading), 0, np.cos(heading)],
|
| 77 |
-
]
|
| 78 |
-
)
|
| 79 |
-
rays = rays @ rot_x.T @ rot_y.T
|
| 80 |
-
|
| 81 |
-
lon = np.arctan2(rays[..., 0], rays[..., 2])
|
| 82 |
-
lat = np.arcsin(np.clip(rays[..., 1], -1.0, 1.0))
|
| 83 |
-
u = ((lon / (2 * np.pi) + 0.5) * src_w).astype(np.int32) % src_w
|
| 84 |
-
v = np.clip(((0.5 - lat / np.pi) * src_h).astype(np.int32), 0, src_h - 1)
|
| 85 |
-
return Image.fromarray(src[v, u])
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
def to_base64(image: Image.Image, fmt: str = "JPEG", quality: int = 85) -> str:
|
| 89 |
-
"""
|
| 90 |
-
Encode an image for transport inside an observation.
|
| 91 |
-
|
| 92 |
-
Args:
|
| 93 |
-
image (`PIL.Image.Image`):
|
| 94 |
-
Image to encode.
|
| 95 |
-
fmt (`str`, *optional*, defaults to `"JPEG"`):
|
| 96 |
-
Pillow format name. Views use JPEG; maps use PNG.
|
| 97 |
-
quality (`int`, *optional*, defaults to `85`):
|
| 98 |
-
JPEG quality, ignored for PNG.
|
| 99 |
-
|
| 100 |
-
Returns:
|
| 101 |
-
`str`: Base64-encoded image bytes, without a data URI prefix.
|
| 102 |
-
"""
|
| 103 |
-
buf = io.BytesIO()
|
| 104 |
-
if fmt.upper() == "JPEG":
|
| 105 |
-
image.save(buf, format="JPEG", quality=quality, optimize=True)
|
| 106 |
-
else:
|
| 107 |
-
image.save(buf, format=fmt)
|
| 108 |
-
return base64.b64encode(buf.getvalue()).decode("ascii")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/build/lib/geoguesser_env/server/scoring.py
DELETED
|
@@ -1,233 +0,0 @@
|
|
| 1 |
-
# SPDX-License-Identifier: BSD-3-Clause
|
| 2 |
-
|
| 3 |
-
"""Distance scoring and action costs.
|
| 4 |
-
|
| 5 |
-
The distance curve is GeoGuessr's own, `5000 * exp(-d / 1492.7)`, normalised
|
| 6 |
-
to `[0, 1]`. Keeping the real curve means scores are directly interpretable
|
| 7 |
-
against the game most people already know.
|
| 8 |
-
"""
|
| 9 |
-
|
| 10 |
-
from __future__ import annotations
|
| 11 |
-
|
| 12 |
-
import math
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
# GeoGuessr's decay constant, in kilometres.
|
| 16 |
-
DECAY_KM = 1492.7
|
| 17 |
-
|
| 18 |
-
# A second, much slower decay used only by the `"mixture"` reward shape. The
|
| 19 |
-
# game curve is worth 0.018 across the whole 6000-20000 km range, so a policy
|
| 20 |
-
# that lands on the wrong continent -- a third of rollouts for a 4B model --
|
| 21 |
-
# gets no gradient for getting *less* wrong. This scale restores one.
|
| 22 |
-
LONG_DECAY_KM = 5000.0
|
| 23 |
-
|
| 24 |
-
# Fraction of the mixture carried by the game curve; the rest is the long scale.
|
| 25 |
-
MIXTURE_SHORT_WEIGHT = 0.5
|
| 26 |
-
|
| 27 |
-
REWARD_SHAPES = ("geoguessr", "mixture")
|
| 28 |
-
|
| 29 |
-
EARTH_RADIUS_KM = 6371.0088
|
| 30 |
-
|
| 31 |
-
# Action costs. Information gathering is cheap but not free, so an episode has
|
| 32 |
-
# to trade breadth of search against committing to a guess.
|
| 33 |
-
COST_LOOK = 0.01
|
| 34 |
-
COST_MAP = 0.01
|
| 35 |
-
COST_PIN = 0.02
|
| 36 |
-
COST_MOVE = 0.05
|
| 37 |
-
|
| 38 |
-
# Partial credit weights, used when the reward is hierarchical.
|
| 39 |
-
WEIGHT_COUNTRY = 0.15
|
| 40 |
-
WEIGHT_REGION = 0.10
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
def haversine_km(lat_a: float, lon_a: float, lat_b: float, lon_b: float) -> float:
|
| 44 |
-
"""
|
| 45 |
-
Great-circle distance between two points in kilometres.
|
| 46 |
-
|
| 47 |
-
Args:
|
| 48 |
-
lat_a (`float`):
|
| 49 |
-
Latitude of the first point in degrees.
|
| 50 |
-
lon_a (`float`):
|
| 51 |
-
Longitude of the first point in degrees.
|
| 52 |
-
lat_b (`float`):
|
| 53 |
-
Latitude of the second point in degrees.
|
| 54 |
-
lon_b (`float`):
|
| 55 |
-
Longitude of the second point in degrees.
|
| 56 |
-
|
| 57 |
-
Returns:
|
| 58 |
-
`float`: Distance in kilometres.
|
| 59 |
-
|
| 60 |
-
Examples:
|
| 61 |
-
|
| 62 |
-
```python
|
| 63 |
-
d = haversine_km(-16.4897, -68.1193, -17.7833, -63.1821)
|
| 64 |
-
```
|
| 65 |
-
"""
|
| 66 |
-
phi_a, phi_b = math.radians(lat_a), math.radians(lat_b)
|
| 67 |
-
d_phi = phi_b - phi_a
|
| 68 |
-
d_lambda = math.radians(lon_b - lon_a)
|
| 69 |
-
h = (
|
| 70 |
-
math.sin(d_phi / 2) ** 2
|
| 71 |
-
+ math.cos(phi_a) * math.cos(phi_b) * math.sin(d_lambda / 2) ** 2
|
| 72 |
-
)
|
| 73 |
-
return 2 * EARTH_RADIUS_KM * math.asin(math.sqrt(min(1.0, h)))
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
def distance_score(distance_km: float, shape: str = "geoguessr") -> float:
|
| 77 |
-
"""
|
| 78 |
-
Map a distance to a score in `[0, 1]`.
|
| 79 |
-
|
| 80 |
-
Args:
|
| 81 |
-
distance_km (`float`):
|
| 82 |
-
Distance between guess and truth in kilometres.
|
| 83 |
-
shape (`str`, *optional*, defaults to `"geoguessr"`):
|
| 84 |
-
`"geoguessr"` for the game's own curve, or `"mixture"` for the
|
| 85 |
-
two-scale curve used when training. The mixture is deliberately
|
| 86 |
-
*not* the default: reported scores stay comparable to the game.
|
| 87 |
-
|
| 88 |
-
Returns:
|
| 89 |
-
`float`: For `"geoguessr"`, `exp(-distance_km / 1492.7)` -- 0 km scores
|
| 90 |
-
`1.0`, 150 km about `0.90`, 5000 km about `0.035`. For `"mixture"`,
|
| 91 |
-
half that plus half of `exp(-distance_km / 5000)`, which keeps the
|
| 92 |
-
mid-range sharp while leaving real gradient past 3000 km.
|
| 93 |
-
|
| 94 |
-
Examples:
|
| 95 |
-
|
| 96 |
-
```python
|
| 97 |
-
near = distance_score(200.0) # 0.875
|
| 98 |
-
far = distance_score(8000.0, shape="mixture") # 0.104, versus 0.005
|
| 99 |
-
```
|
| 100 |
-
"""
|
| 101 |
-
if distance_km < 0:
|
| 102 |
-
raise ValueError(f"distance_km must be non-negative, got {distance_km}")
|
| 103 |
-
if shape not in REWARD_SHAPES:
|
| 104 |
-
raise ValueError(f"shape must be one of {REWARD_SHAPES}, got {shape!r}")
|
| 105 |
-
short = math.exp(-distance_km / DECAY_KM)
|
| 106 |
-
if shape == "geoguessr":
|
| 107 |
-
return short
|
| 108 |
-
long = math.exp(-distance_km / LONG_DECAY_KM)
|
| 109 |
-
return MIXTURE_SHORT_WEIGHT * short + (1.0 - MIXTURE_SHORT_WEIGHT) * long
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
def action_cost(
|
| 113 |
-
n_looks: int = 0, n_maps: int = 0, n_pins: int = 0, n_moves: int = 0
|
| 114 |
-
) -> float:
|
| 115 |
-
"""
|
| 116 |
-
Total cost of the information gathering done this episode.
|
| 117 |
-
|
| 118 |
-
Args:
|
| 119 |
-
n_looks (`int`, *optional*, defaults to `0`):
|
| 120 |
-
Number of view renders.
|
| 121 |
-
n_maps (`int`, *optional*, defaults to `0`):
|
| 122 |
-
Number of map views that were not pins.
|
| 123 |
-
n_pins (`int`, *optional*, defaults to `0`):
|
| 124 |
-
Number of pins placed.
|
| 125 |
-
n_moves (`int`, *optional*, defaults to `0`):
|
| 126 |
-
Number of moves taken.
|
| 127 |
-
|
| 128 |
-
Returns:
|
| 129 |
-
`float`: Cost to subtract from the distance score.
|
| 130 |
-
"""
|
| 131 |
-
return (
|
| 132 |
-
COST_LOOK * n_looks
|
| 133 |
-
+ COST_MAP * n_maps
|
| 134 |
-
+ COST_PIN * n_pins
|
| 135 |
-
+ COST_MOVE * n_moves
|
| 136 |
-
)
|
| 137 |
-
|
| 138 |
-
|
| 139 |
-
# The multiplicative cost is capped so a very long episode scales the score down
|
| 140 |
-
# rather than erasing it. Without a cap a 20-move episode would reach zero.
|
| 141 |
-
MAX_COST_FRACTION = 0.5
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
def compute_reward(
|
| 145 |
-
distance_km: float | None,
|
| 146 |
-
*,
|
| 147 |
-
cost: float = 0.0,
|
| 148 |
-
country_hit: bool = False,
|
| 149 |
-
region_hit: bool = False,
|
| 150 |
-
hierarchical: bool = False,
|
| 151 |
-
shape: str = "geoguessr",
|
| 152 |
-
cost_mode: str = "subtract",
|
| 153 |
-
) -> float:
|
| 154 |
-
"""
|
| 155 |
-
Combine distance, partial credit and action cost into one reward.
|
| 156 |
-
|
| 157 |
-
Args:
|
| 158 |
-
distance_km (`float` or `None`):
|
| 159 |
-
Distance from guess to truth. `None` means the guess could not be
|
| 160 |
-
parsed, which scores zero before costs.
|
| 161 |
-
cost (`float`, *optional*, defaults to `0.0`):
|
| 162 |
-
Accumulated action cost.
|
| 163 |
-
country_hit (`bool`, *optional*, defaults to `False`):
|
| 164 |
-
Whether the guessed country matched.
|
| 165 |
-
region_hit (`bool`, *optional*, defaults to `False`):
|
| 166 |
-
Whether the guessed region matched.
|
| 167 |
-
hierarchical (`bool`, *optional*, defaults to `False`):
|
| 168 |
-
Whether to add country and region partial credit.
|
| 169 |
-
shape (`str`, *optional*, defaults to `"geoguessr"`):
|
| 170 |
-
Distance curve, passed to [`~scoring.distance_score`].
|
| 171 |
-
cost_mode (`str`, *optional*, defaults to `"subtract"`):
|
| 172 |
-
`"subtract"` reproduces the game: the cost comes off the score and
|
| 173 |
-
the result is floored at zero. `"multiply"` scales the score by
|
| 174 |
-
`1 - cost` instead, which is what training wants -- see below.
|
| 175 |
-
|
| 176 |
-
Returns:
|
| 177 |
-
`float`: Reward in `[0, 1]`.
|
| 178 |
-
|
| 179 |
-
<Tip warning={true}>
|
| 180 |
-
|
| 181 |
-
`"subtract"` and a floor at zero destroy the ordering of bad guesses. Mean
|
| 182 |
-
cost for a 4B model is 0.13, and the game curve falls below that at about
|
| 183 |
-
3300 km, so a 3324 km miss and an 18723 km miss both score exactly 0.0 --
|
| 184 |
-
measured across 200 episodes, 77 of them collapsed to a single value with
|
| 185 |
-
zero variance. A GRPO group drawn from those has no advantage and therefore
|
| 186 |
-
contributes no gradient. `"multiply"` cannot do this: scaling by a positive
|
| 187 |
-
factor preserves the ordering whatever the cost.
|
| 188 |
-
|
| 189 |
-
</Tip>
|
| 190 |
-
|
| 191 |
-
Examples:
|
| 192 |
-
|
| 193 |
-
```python
|
| 194 |
-
# Training: gradient survives on the wrong continent.
|
| 195 |
-
reward = compute_reward(8000.0, cost=0.13, shape="mixture", cost_mode="multiply")
|
| 196 |
-
```
|
| 197 |
-
"""
|
| 198 |
-
if distance_km is None:
|
| 199 |
-
return 0.0
|
| 200 |
-
score = distance_score(distance_km, shape=shape)
|
| 201 |
-
if hierarchical:
|
| 202 |
-
score += WEIGHT_COUNTRY * country_hit + WEIGHT_REGION * region_hit
|
| 203 |
-
score = min(1.0, score)
|
| 204 |
-
if cost_mode == "multiply":
|
| 205 |
-
return score * (1.0 - min(max(cost, 0.0), MAX_COST_FRACTION))
|
| 206 |
-
if cost_mode != "subtract":
|
| 207 |
-
raise ValueError(
|
| 208 |
-
f"cost_mode must be 'subtract' or 'multiply', got {cost_mode!r}"
|
| 209 |
-
)
|
| 210 |
-
return max(0.0, score - cost)
|
| 211 |
-
|
| 212 |
-
|
| 213 |
-
def verdict(distance_km: float) -> str:
|
| 214 |
-
"""
|
| 215 |
-
Short human-readable label for a distance, for UIs and logs.
|
| 216 |
-
|
| 217 |
-
Args:
|
| 218 |
-
distance_km (`float`):
|
| 219 |
-
Distance between guess and truth in kilometres.
|
| 220 |
-
|
| 221 |
-
Returns:
|
| 222 |
-
`str`: One of `"Perfect"`, `"Pinpoint"`, `"Close"`, `"Right region"`
|
| 223 |
-
or `"Wrong continent"`.
|
| 224 |
-
"""
|
| 225 |
-
if distance_km < 0.025:
|
| 226 |
-
return "Perfect"
|
| 227 |
-
if distance_km < 25:
|
| 228 |
-
return "Pinpoint"
|
| 229 |
-
if distance_km < 200:
|
| 230 |
-
return "Close"
|
| 231 |
-
if distance_km < 1500:
|
| 232 |
-
return "Right region"
|
| 233 |
-
return "Wrong continent"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/openenv_geoguesser_env.egg-info/PKG-INFO
DELETED
|
@@ -1,760 +0,0 @@
|
|
| 1 |
-
Metadata-Version: 2.4
|
| 2 |
-
Name: openenv-geoguesser-env
|
| 3 |
-
Version: 0.1.0
|
| 4 |
-
Summary: GeoGuessr-style visual geolocation environment for OpenEnv
|
| 5 |
-
Requires-Python: >=3.10
|
| 6 |
-
Description-Content-Type: text/markdown
|
| 7 |
-
Requires-Dist: openenv>=0.3.1
|
| 8 |
-
Requires-Dist: fastapi>=0.115.0
|
| 9 |
-
Requires-Dist: pydantic>=2.0.0
|
| 10 |
-
Requires-Dist: uvicorn>=0.24.0
|
| 11 |
-
Requires-Dist: fastmcp>=2.0.0
|
| 12 |
-
Requires-Dist: pillow>=10.0.0
|
| 13 |
-
Requires-Dist: numpy>=1.24.0
|
| 14 |
-
Requires-Dist: matplotlib>=3.7.0
|
| 15 |
-
Provides-Extra: ui
|
| 16 |
-
Requires-Dist: gradio>=4.0.0; extra == "ui"
|
| 17 |
-
Provides-Extra: dev
|
| 18 |
-
Requires-Dist: pytest>=8.0.0; extra == "dev"
|
| 19 |
-
|
| 20 |
-
# GeoGuesser
|
| 21 |
-
|
| 22 |
-
A GeoGuessr-style visual geolocation environment. The agent is dropped at an
|
| 23 |
-
unknown street-level location, looks around, walks along the road, pins
|
| 24 |
-
candidate coordinates on a map to check itself, and commits to a final guess.
|
| 25 |
-
Reward is distance-based, using the game's own scoring curve.
|
| 26 |
-
|
| 27 |
-
Independent open-source project, unaffiliated with GeoGuessr AB. Imagery comes
|
| 28 |
-
from Mapillary contributors under CC-BY-SA-4.0.
|
| 29 |
-
|
| 30 |
-
## Quick start
|
| 31 |
-
|
| 32 |
-
```bash
|
| 33 |
-
cd envs/geoguesser_env
|
| 34 |
-
|
| 35 |
-
# The frozen 200-task eval split is committed, so this runs as-is, with the
|
| 36 |
-
# same configuration the Space uses.
|
| 37 |
-
./scripts/serve_local.sh # http://localhost:8000/web/
|
| 38 |
-
|
| 39 |
-
# To build your own data (needs a free Mapillary token with READ scope):
|
| 40 |
-
export MAPILLARY_API_KEY_TRAIN="MLY|..."
|
| 41 |
-
python scripts/harvest_tiles.py # enumerate sequences
|
| 42 |
-
./scripts/build_dataset.sh # mirror tasks offline
|
| 43 |
-
python scripts/verify_offline.py tasks/pool_offline_5k.jsonl
|
| 44 |
-
python scripts/split_tasks.py tasks/pool_offline_5k.jsonl --eval 200
|
| 45 |
-
```
|
| 46 |
-
|
| 47 |
-
```python
|
| 48 |
-
from geoguesser_env import GeoGuesserEnv, GuessAction, LookAction, PinAction
|
| 49 |
-
|
| 50 |
-
env = GeoGuesserEnv(base_url="http://localhost:8000")
|
| 51 |
-
|
| 52 |
-
result = env.reset(split="eval", index=7) # byte-identical on repeat
|
| 53 |
-
print(result.observation.prompt)
|
| 54 |
-
|
| 55 |
-
result = env.step(LookAction(heading_deg=90, fov_deg=45))
|
| 56 |
-
result = env.step(PinAction(lat=-16.5, lon=-68.1))
|
| 57 |
-
print(result.observation.feedback)
|
| 58 |
-
# Pin 1 placed at -16.5000, -68.1000 - Bolivia (South America).
|
| 59 |
-
# Nearest major city: La Paz, ~5 km E. 10 actions left.
|
| 60 |
-
|
| 61 |
-
result = env.step(GuessAction(response="Altiplano. <guess>-16.49, -68.12</guess>"))
|
| 62 |
-
print(result.reward, result.observation.distance_km)
|
| 63 |
-
```
|
| 64 |
-
|
| 65 |
-
## Tools
|
| 66 |
-
|
| 67 |
-
| Tool | What it does | Cost |
|
| 68 |
-
|------|--------------|------|
|
| 69 |
-
| `look(heading_deg, pitch_deg, fov_deg)` | Render a view. Heading is absolute, `0` is true north | −0.01 |
|
| 70 |
-
| `pan(delta_deg)` | Turn relative to the current heading | −0.01 |
|
| 71 |
-
| `zoom(fov_deg)` | Narrow the field of view; around 30 reads distant signs | −0.01 |
|
| 72 |
-
| `move(direction, meters)` | Walk the captured road; reports distance actually travelled | −0.05 |
|
| 73 |
-
| `place_pin(lat, lon, label)` | Pin a candidate and see where it falls on the map | −0.02 |
|
| 74 |
-
| `view_map(lat, lon, span_deg)` | Pan and zoom the map without pinning | −0.01 |
|
| 75 |
-
| `list_pins()` / `clear_pins()` | Review or drop candidates | free |
|
| 76 |
-
| `measure(lat_a, lon_a, lat_b, lon_b)` | Distance between two of your own points | free |
|
| 77 |
-
| `reverse_geocode(lat, lon)` | Name the country and nearest city at a coordinate | free |
|
| 78 |
-
| `submit_guess(lat, lon, ...)` | Commit the answer. Terminal | — |
|
| 79 |
-
|
| 80 |
-
Tools the backend cannot serve are **not registered**, so the agent never sees
|
| 81 |
-
a tool that always fails.
|
| 82 |
-
|
| 83 |
-
### The agent's map and the player's map agree
|
| 84 |
-
|
| 85 |
-
The player sees live OpenFreeMap tiles; the agent sees an offline Natural Earth
|
| 86 |
-
render. They have to agree about *how precisely a pin can be aimed*, because
|
| 87 |
-
that is what the distance reward measures — a map showing only country outlines
|
| 88 |
-
lets you place a country, not a point within a city.
|
| 89 |
-
|
| 90 |
-
So the guess map is zoom-aware. `place_pin` takes `span_deg`, and the render
|
| 91 |
-
adds detail as the window tightens:
|
| 92 |
-
|
| 93 |
-
| Window | What the agent's map shows |
|
| 94 |
-
|--------|----------------------------|
|
| 95 |
-
| wider than ~4 deg | coastlines, borders, country names |
|
| 96 |
-
| under ~4 deg | urban areas, highways, rivers, town names (Natural Earth 10m) |
|
| 97 |
-
| under ~0.35 deg | **real OSM streets**, fetched from Overpass and cached |
|
| 98 |
-
|
| 99 |
-
Natural Earth tops out at highway level — it shows the motorways around a city
|
| 100 |
-
but not the grid inside it. Below 0.35 degrees the map therefore fetches actual
|
| 101 |
-
ways from Overpass, generalising by zoom the way a real style does: minor
|
| 102 |
-
classes appear only once the window is tight enough to hold them, and widths
|
| 103 |
-
grow as it shrinks. A pin on Abuja at `span_deg=0.05` came back as an 11 km
|
| 104 |
-
window with the full street grid, drawn white-on-pale to read like the
|
| 105 |
-
player's Positron tiles.
|
| 106 |
-
|
| 107 |
-
Overpass has real limits, and they are the binding constraint on how this
|
| 108 |
-
scales: roughly **10,000 requests and 1 GB per day**, about **2 concurrent
|
| 109 |
-
slots per IP**, a 180 s runtime and 512 MiB memory ceiling per query, HTTP 429
|
| 110 |
-
when rate limited and 504 when a query is too large. Cooldowns lengthen for
|
| 111 |
-
heavy users. So street detail is right for eval, demos and modest training, and
|
| 112 |
-
the cache is what keeps it polite — a run doing millions of pins must pre-warm
|
| 113 |
-
or bundle a Protomaps extract instead.
|
| 114 |
-
|
| 115 |
-
The first render of a neighbourhood costs 3-16 s; every later one is served
|
| 116 |
-
from `data/geo/osm_cache/` in ~30 ms and is byte-identical. That makes an
|
| 117 |
-
episode deterministic once warm, and a frozen eval should pre-warm the cache
|
| 118 |
-
the same way it pre-warms panoramas — or set
|
| 119 |
-
`GEOGUESSER_STREET_DETAIL=0`, which falls back to Natural Earth and never
|
| 120 |
-
touches the network. Any fetch failure degrades to no streets rather than
|
| 121 |
-
failing the step.
|
| 122 |
-
|
| 123 |
-
The optional detail layers are fetched once, since 87 MB of GeoJSON does not
|
| 124 |
-
belong in the repo:
|
| 125 |
-
|
| 126 |
-
```bash
|
| 127 |
-
python scripts/fetch_detail_geo.py # compacts to ~39 MB, gitignored
|
| 128 |
-
```
|
| 129 |
-
|
| 130 |
-
Without them the map still renders, with outlines and major cities only.
|
| 131 |
-
|
| 132 |
-
In the play page the pin carries the zoom you are actually looking at, so the
|
| 133 |
-
"what the agent sees" panel is framed like your own view at the same scale and
|
| 134 |
-
with comparable detail.
|
| 135 |
-
|
| 136 |
-
Overpass has a usage policy that discourages heavy automated querying, so this
|
| 137 |
-
is right for eval, demos and modest training, and the cache is what keeps it
|
| 138 |
-
polite. A run doing millions of pins should pre-warm or bundle a Protomaps
|
| 139 |
-
extract instead.
|
| 140 |
-
|
| 141 |
-
### Zoom resolves real detail
|
| 142 |
-
|
| 143 |
-
Zooming is not cosmetic, but it needs the right source. A 30-degree view of a
|
| 144 |
-
2048x1024 panorama samples only about 170 source pixels, so narrowing the field
|
| 145 |
-
of view barely adds information — measured mean gradient 6.60 at 90 degrees
|
| 146 |
-
against 7.03 at 30. The 7680x3840 original roughly doubles it (10.23 against
|
| 147 |
-
14.87), which is the difference between guessing at a sign and reading it.
|
| 148 |
-
|
| 149 |
-
So each panorama is cached twice. Wide views render from the 2048 derivative in
|
| 150 |
-
~30 ms; a field of view at or below 45 degrees pulls the original and renders in
|
| 151 |
-
~70 ms. If no original exists the step degrades to a soft view rather than
|
| 152 |
-
failing. Set `GEOGUESSER_HIRES_ZOOM=0` to disable it.
|
| 153 |
-
|
| 154 |
-
### Pinning tells you where you pointed, not whether you are right
|
| 155 |
-
|
| 156 |
-
`place_pin` returns a rendered map and a description of the pinned location:
|
| 157 |
-
country, subregion, nearest city with distance and bearing, and the distance
|
| 158 |
-
to the agent's own earlier pins. It reveals nothing about the target.
|
| 159 |
-
|
| 160 |
-
That restraint is deliberate. Any signal about the truth — a distance, a
|
| 161 |
-
warmer/colder hint — would make binary search the optimal policy, and the
|
| 162 |
-
environment would measure bisection rather than geographic reasoning. Distance
|
| 163 |
-
and score arrive only from `submit_guess`.
|
| 164 |
-
|
| 165 |
-
## Reward
|
| 166 |
-
|
| 167 |
-
```
|
| 168 |
-
geo = exp(-distance_km / 1492.7) # GeoGuessr's curve, in [0, 1]
|
| 169 |
-
partial = 0.15 * country_hit + 0.10 * region_hit # when hierarchical
|
| 170 |
-
cost = 0.01*looks + 0.01*maps + 0.02*pins + 0.05*moves
|
| 171 |
-
reward = clip(geo + partial, 0, 1) - cost
|
| 172 |
-
```
|
| 173 |
-
|
| 174 |
-
An unparseable or out-of-range guess scores `0.0` and says why. Parsing
|
| 175 |
-
accepts what models actually emit: decimal pairs, DMS (`48°51'29"N`), labelled
|
| 176 |
-
`lat:`/`lon:`, JSON, and `<guess>` tags.
|
| 177 |
-
|
| 178 |
-
### Training against it
|
| 179 |
-
|
| 180 |
-
The defaults above reproduce the game, which makes a score directly comparable
|
| 181 |
-
to GeoGuessr. They are the wrong shape for RL, and three flags change that:
|
| 182 |
-
|
| 183 |
-
| flag | play / eval | training | why |
|
| 184 |
-
|---|---|---|---|
|
| 185 |
-
| `reward_shape` | `"geoguessr"` | `"mixture"` | The game curve is worth **0.018** across the whole 6000-20000 km range, so a policy gets no gradient for landing on the right continent instead of the wrong one. `"mixture"` adds a 5000 km scale, making that span worth 0.150. |
|
| 186 |
-
| `cost_mode` | `"subtract"` | `"multiply"` | Mean action cost for a 4B model is 0.13 and the curve falls below that at ~3300 km, so `max(0, geo - cost)` floors every worse guess at exactly zero. Measured over 200 episodes: **77 collapsed to 0.0 with zero variance**, so a GRPO group drawn from them has no advantage and yields no gradient. A multiplier cannot do this. |
|
| 187 |
-
| `hide_task_identity` | `False` | **`True`** | `metadata` carries `attribution.creator_username`, and the contributor determines the country outright for **74% of training tasks** (`amsterdam` only maps the Netherlands). `task_index`/`task_id`/`sequence_id` are a few thousand memorisable keys straight to a coordinate. Either lets a policy score without reading the image. |
|
| 188 |
-
|
| 189 |
-
```bash
|
| 190 |
-
GEOGUESSER_REWARD_SHAPE=mixture \
|
| 191 |
-
GEOGUESSER_COST_MODE=multiply \
|
| 192 |
-
GEOGUESSER_HIDE_IDENTITY=1 \
|
| 193 |
-
uvicorn geoguesser_env.server.app:app
|
| 194 |
-
```
|
| 195 |
-
|
| 196 |
-
Replaying all 3,037 recorded eval episodes through both settings: episodes
|
| 197 |
-
scoring exactly zero fall from **10-46% to 0%** for every model, and the
|
| 198 |
-
leaderboard order only changes within the tiers already documented as inside
|
| 199 |
-
noise at n=200.
|
| 200 |
-
|
| 201 |
-
The terminal observation carries full provenance either way — once the truth is
|
| 202 |
-
revealed it can no longer be used to shortcut the episode — so recorded traces
|
| 203 |
-
stay complete under `hide_task_identity`.
|
| 204 |
-
|
| 205 |
-
## Baselines
|
| 206 |
-
|
| 207 |
-
Measured with `examples/geoguesser_llm_rollout.py` on the committed index,
|
| 208 |
-
tasks 0/7/14/21/28, so the numbers are reproducible rather than illustrative.
|
| 209 |
-
Five episodes is far too few for a leaderboard; they are a smoke test that the
|
| 210 |
-
task is solvable and the reward is discriminative.
|
| 211 |
-
|
| 212 |
-
| Model | Mode | Mean reward | Median distance | Within 200 km | Parsed |
|
| 213 |
-
|-------|------|------------:|----------------:|--------------:|-------:|
|
| 214 |
-
| `claude-sonnet-5` | single-shot | **0.896** | 98 km | 4/5 | 5/5 |
|
| 215 |
-
| `claude-sonnet-5` | agentic, tasks 0-5 | 0.539-0.653 | 574-660 km | 3/6 | 6/6 |
|
| 216 |
-
| `Qwen/Qwen3.5-9B` | agentic, tasks 0-5 | 0.355 | 1,136 km | 0/5 | 5/6 |
|
| 217 |
-
| `Qwen/Qwen3.5-9B:together` | single-shot | 0.277 | 1,139 km | 1/3 | 3/5 |
|
| 218 |
-
| `Qwen/Qwen3.5-9B:together` | agentic, 6 turns | 0.304 | 579 km | 0/1 | 1/2 |
|
| 219 |
-
| `Qwen/Qwen3.5-9B:together` | agentic, 8k tokens | 0.087 | 2,423 km | 0/1 | **4/4 turns** |
|
| 220 |
-
| `Qwen/Qwen3.5-9B:together` | single-shot, 8 eps, 4 parallel | 0.412 | 787 km | 1/6 | 6/8 |
|
| 221 |
-
|
| 222 |
-
Sonnet placed two guesses within 2 km. The agentic score sits slightly below
|
| 223 |
-
single-shot on the same tasks because looking around costs reward and the extra
|
| 224 |
-
views did not always pay for themselves — which is the trade-off the
|
| 225 |
-
environment is meant to expose, not a defect.
|
| 226 |
-
|
| 227 |
-
Qwen does follow the multi-turn protocol: across the agentic runs it produced
|
| 228 |
-
`look`, `move`, `zoom`, `pin` and `guess` actions and navigated up to 68 m down
|
| 229 |
-
a road. Two things had to be right first, and both are prompting or plumbing
|
| 230 |
-
rather than capability:
|
| 231 |
-
|
| 232 |
-
- **Token budget.** Reasoning models put their chain of thought in a separate
|
| 233 |
-
`reasoning_content` field and can exhaust the budget before emitting any
|
| 234 |
-
content, which looks exactly like a model that cannot see images. At 1,024
|
| 235 |
-
tokens Qwen scored 0/5 with empty replies; at 3,500 it followed the protocol
|
| 236 |
-
intermittently, failing turns whose reply came back as pure reasoning; at
|
| 237 |
-
8,000 it parsed 4/4 turns. The example defaults to 3,000 and takes
|
| 238 |
-
`--max-tokens`.
|
| 239 |
-
- **A deadline.** Left to itself Qwen explored until the turn budget ran out and
|
| 240 |
-
scored zero. The agentic prompt now warns it explicitly when two turns remain,
|
| 241 |
-
after which it committed.
|
| 242 |
-
|
| 243 |
-
What remains is accuracy, not plumbing: its guesses landed 452 km, 579 km and
|
| 244 |
-
2,423 km out against Sonnet's 98 km median. It is also 10x slower — 110-193 s
|
| 245 |
-
per agentic episode against Sonnet's 10-16 s.
|
| 246 |
-
|
| 247 |
-
```bash
|
| 248 |
-
python examples/geoguesser_llm_rollout.py --provider anthropic \
|
| 249 |
-
--model claude-sonnet-5 --episodes 5
|
| 250 |
-
python examples/geoguesser_llm_rollout.py --provider hf \
|
| 251 |
-
--model "Qwen/Qwen3.5-9B:together" --episodes 5 --max-tokens 4000
|
| 252 |
-
python examples/geoguesser_llm_rollout.py --provider anthropic \
|
| 253 |
-
--mode agentic --episodes 3 --verbose
|
| 254 |
-
```
|
| 255 |
-
|
| 256 |
-
## Parallel rollouts
|
| 257 |
-
|
| 258 |
-
An episode is stateful, so concurrent rollouts each need their own environment
|
| 259 |
-
instance over the shared read-only index and cache. `--concurrency` does that:
|
| 260 |
-
|
| 261 |
-
```bash
|
| 262 |
-
python examples/geoguesser_llm_rollout.py --provider hf \
|
| 263 |
-
--model "Qwen/Qwen3.5-9B:together" --episodes 8 --concurrency 4
|
| 264 |
-
```
|
| 265 |
-
|
| 266 |
-
Eight Qwen episodes took 113.6 s wall against 334.0 s of summed latency — a
|
| 267 |
-
**2.94x speedup** on 4 workers, the shortfall being the provider's own queuing
|
| 268 |
-
rather than the environment, which spends ~28 ms on a reset.
|
| 269 |
-
|
| 270 |
-
## Scaling: what the environment can actually supply
|
| 271 |
-
|
| 272 |
-
Measured on an 18-core machine with a warm cache and street detail off, one
|
| 273 |
-
environment per worker, with a correctness assertion in the loop so an
|
| 274 |
-
interference bug cannot masquerade as throughput.
|
| 275 |
-
|
| 276 |
-
Per-step cost is dominated by map rendering, not imagery:
|
| 277 |
-
|
| 278 |
-
| Step | Cost |
|
| 279 |
-
|------|-----:|
|
| 280 |
-
| reset, or `look` at 90 deg fov | **29 ms** |
|
| 281 |
-
| `look` at 30 deg fov, from the original | 74 ms |
|
| 282 |
-
| `place_pin`, a two-panel map | **207 ms** |
|
| 283 |
-
| `submit_guess` with the reveal map | 278 ms |
|
| 284 |
-
|
| 285 |
-
That shapes the two configurations:
|
| 286 |
-
|
| 287 |
-
| Config | Workers | Episodes/s | Notes |
|
| 288 |
-
|--------|--------:|-----------:|-------|
|
| 289 |
-
| training — views only, no pins, no reveal | 8 threads | **31.7** | 158 env steps/s |
|
| 290 |
-
| eval — pins and reveal map | 4 processes | **3.5** | matplotlib is GIL-bound |
|
| 291 |
-
| eval — pins and reveal map | 8 threads | 1.9 | threads do not help here |
|
| 292 |
-
|
| 293 |
-
Two facts fall out of that. `look` releases the GIL — the reprojection is numpy
|
| 294 |
-
and the encode is Pillow — so **threads scale well for view-only work** and
|
| 295 |
-
processes only add startup cost. Map rendering is pure Python, so it is
|
| 296 |
-
**GIL-bound and needs processes**, which buy about 1.8x before contention.
|
| 297 |
-
|
| 298 |
-
`GEOGUESSER_REVEAL_MAP=0` is the single biggest throughput lever: a
|
| 299 |
-
training-shaped episode drops from **389 ms to 117 ms**, since the reveal map is
|
| 300 |
-
280 ms that a training run never reads — the reward and the distance are in the
|
| 301 |
-
observation either way. Keep it on for evals, traces and the UI.
|
| 302 |
-
|
| 303 |
-
Memory is small: about 215 MB for the imports, 100 MB more once the geodata
|
| 304 |
-
caches fill, and **~13 MB per additional environment in the same process**. The
|
| 305 |
-
Natural Earth layers are process-shared through an LRU cache, so threads are far
|
| 306 |
-
cheaper than processes here too.
|
| 307 |
-
|
| 308 |
-
### Rate limits that actually bind
|
| 309 |
-
|
| 310 |
-
| API | When it is called | Limit |
|
| 311 |
-
|-----|-------------------|-------|
|
| 312 |
-
| Mapillary | only on a cache miss | 60,000/min entity, 10,000/min search, 50,000/day tiles |
|
| 313 |
-
| Overpass | only a pin below 0.35 deg span | ~10,000/day, 2 concurrent slots per IP |
|
| 314 |
-
| the model | every turn | the real constraint |
|
| 315 |
-
|
| 316 |
-
With a warm cache the environment makes **no network calls at all**. Default
|
| 317 |
-
pins use a 7 degree span, which is above the street threshold, so they do not
|
| 318 |
-
touch Overpass either — only a deliberately zoomed pin does. For an eval sweep
|
| 319 |
-
set `GEOGUESSER_STREET_DETAIL=0` unless you have pre-warmed, since 100 zoomed
|
| 320 |
-
pins against 2 concurrent slots would throttle immediately.
|
| 321 |
-
|
| 322 |
-
### 100-task eval
|
| 323 |
-
|
| 324 |
-
Sonnet agentic measured 13.1 s per episode, so 100 episodes cost about 1,310
|
| 325 |
-
model-seconds and the concurrency you can use is set by the provider, not by
|
| 326 |
-
this environment:
|
| 327 |
-
|
| 328 |
-
- concurrency 4 → **~5.5 minutes**
|
| 329 |
-
- concurrency 8 → **~2.7 minutes**, if your token-per-minute ceiling allows it
|
| 330 |
-
|
| 331 |
-
An agentic episode sends roughly five 640x640 views, about 550 tokens each, plus
|
| 332 |
-
a growing text prompt. At a 400k input-tokens-per-minute ceiling that is roughly
|
| 333 |
-
6 episodes in flight before tokens, not latency, become the limit — which is why
|
| 334 |
-
4 to 6 is the practical range for Sonnet. Qwen through the router reached 3.97x
|
| 335 |
-
on 6 workers, so 6 to 8 there.
|
| 336 |
-
|
| 337 |
-
Meanwhile the environment can supply 3.5 episodes/s in eval configuration
|
| 338 |
-
against the 0.3 to 0.6 episodes/s those concurrencies actually consume, so it
|
| 339 |
-
has roughly ten times the headroom it needs.
|
| 340 |
-
|
| 341 |
-
### 1000 training steps
|
| 342 |
-
|
| 343 |
-
For GRPO with 8 prompts and a group of 16, that is 128 episodes per step and
|
| 344 |
-
**128,000 episodes**, or about 640,000 env steps:
|
| 345 |
-
|
| 346 |
-
- environment time: 128,000 / 31.7 ≈ **1.1 hours total**, spread across workers
|
| 347 |
-
- imagery: nothing, with a warm cache
|
| 348 |
-
- storage: 100 tasks warm is 60 MB; 5,000 tasks would be about 1.5 GB of start
|
| 349 |
-
frames
|
| 350 |
-
|
| 351 |
-
So the environment is not the constraint — 640,000 model calls are, which needs
|
| 352 |
-
batched local inference rather than an API. The constraint that *is* ours is
|
| 353 |
-
**task diversity**: 128,000 episodes over 100 tasks means each location is seen
|
| 354 |
-
1,280 times, which is memorisation territory. Before a run that long, either
|
| 355 |
-
harvest more tasks or add seeded heading augmentation, which multiplies
|
| 356 |
-
effective tasks 8 to 12 times from imagery already on disk.
|
| 357 |
-
|
| 358 |
-
## Readiness audit
|
| 359 |
-
|
| 360 |
-
`scripts/readiness_check.py` checks the properties that only appear at real
|
| 361 |
-
scale and concurrency, rather than on the four committed fixtures:
|
| 362 |
-
|
| 363 |
-
```bash
|
| 364 |
-
python scripts/readiness_check.py --full
|
| 365 |
-
```
|
| 366 |
-
|
| 367 |
-
```
|
| 368 |
-
[PASS] index integrity 100 tasks, 47 countries, 100 unique sequences
|
| 369 |
-
[PASS] all tasks render 100/100 rendered, median reset 28 ms
|
| 370 |
-
[PASS] cross-process determinism two subprocesses and this process agree
|
| 371 |
-
[PASS] parallel isolation 8 concurrent episodes, each its own task
|
| 372 |
-
[PASS] offline with warm cache 12/12 served with fetching disabled
|
| 373 |
-
[PASS] reward is discriminative uniform-random 0.029, fixed-point 0.111
|
| 374 |
-
[PASS] step latency look 29 ms, pin + map 249 ms
|
| 375 |
-
[PASS] one guess per episode a second guess returns reward=None
|
| 376 |
-
[PASS] pin never leaks the target 60 pins across 20 tasks revealed nothing
|
| 377 |
-
```
|
| 378 |
-
|
| 379 |
-
The reward check matters most: a uniform-random guesser scores **0.029** and the
|
| 380 |
-
best trivial constant guess **0.111**, against Sonnet's 0.896. The signal is
|
| 381 |
-
measuring geolocation rather than rewarding noise.
|
| 382 |
-
|
| 383 |
-
## Tracing a rollout
|
| 384 |
-
|
| 385 |
-
`--trace-dir` records every turn: the image the model saw, what it said, the
|
| 386 |
-
action it chose, the environment's reply, steps left and running cost. A second
|
| 387 |
-
script renders that as one self-contained HTML page, which is the difference
|
| 388 |
-
between knowing the reward and seeing why:
|
| 389 |
-
|
| 390 |
-
```bash
|
| 391 |
-
python examples/geoguesser_llm_rollout.py --provider anthropic \
|
| 392 |
-
--mode agentic --episodes 6 --trace-dir rollouts
|
| 393 |
-
python scripts/render_trace.py rollouts/anthropic_agentic
|
| 394 |
-
```
|
| 395 |
-
|
| 396 |
-
Images are written beside the trace rather than inlined, since six agentic
|
| 397 |
-
episodes carry around 35 views and a JSONL with those in it is neither readable
|
| 398 |
-
nor loadable.
|
| 399 |
-
|
| 400 |
-
Every step carries an image, including the guess: a guess returns a **reveal
|
| 401 |
-
map** with the guess, the true location and the line between them. Truth is
|
| 402 |
-
drawn only there, after scoring. When the guess is more than 25 degrees out the
|
| 403 |
-
second panel frames the true location instead of both points, because squashing
|
| 404 |
-
a hemisphere into a panel shows nothing.
|
| 405 |
-
|
| 406 |
-
## On the Hub
|
| 407 |
-
|
| 408 |
-
| | |
|
| 409 |
-
|---|---|
|
| 410 |
-
| Space | [`HuggingEnvs/geoguesser-env`](https://huggingface.co/spaces/HuggingEnvs/geoguesser-env) |
|
| 411 |
-
| Space (mirror) | [`AdithyaSK/geoguesser-env`](https://huggingface.co/spaces/AdithyaSK/geoguesser-env) |
|
| 412 |
-
| Task splits | [`HuggingEnvs/geoguesser-tasks`](https://huggingface.co/datasets/HuggingEnvs/geoguesser-tasks) |
|
| 413 |
-
| Imagery | [`HuggingEnvs/geoguesser-panos`](https://huggingface.co/buckets/HuggingEnvs/geoguesser-panos) (Storage Bucket, public, mounted read-only at `/data`) |
|
| 414 |
-
|
| 415 |
-
The Space and the bucket deliberately live in different namespaces: moving the
|
| 416 |
-
Space should not mean re-uploading 22 GB, so `deploy_hub.py` takes the owner of
|
| 417 |
-
each separately (`GEOGUESSER_HF_ORG` and `GEOGUESSER_HF_BUCKET_OWNER`).
|
| 418 |
-
|
| 419 |
-
The same client drives either:
|
| 420 |
-
|
| 421 |
-
```python
|
| 422 |
-
env = GeoGuesserEnv(base_url="http://localhost:8000") # local
|
| 423 |
-
env = GeoGuesserEnv(base_url="https://huggingenvs-geoguesser-env.hf.space") # Space
|
| 424 |
-
```
|
| 425 |
-
|
| 426 |
-
Only the file locations differ — locally the indexes and imagery are in the
|
| 427 |
-
repo, on the Space they arrive through the bucket mount. Splits, step budget,
|
| 428 |
-
street labels and offline enforcement are identical, and verified so: the same
|
| 429 |
-
task returns the same image checksums, reward and distance from both.
|
| 430 |
-
|
| 431 |
-
Deploy with `scripts/deploy_hub.py --all`. `openenv push` is deliberately not
|
| 432 |
-
used: it cannot attach a bucket volume, and its default excludes would upload
|
| 433 |
-
22 GB of panoramas into git.
|
| 434 |
-
|
| 435 |
-
### Local and Space parity
|
| 436 |
-
|
| 437 |
-
Verified across a local server and both Spaces, on both splits, driven by the
|
| 438 |
-
same client. Every field below is identical in all three:
|
| 439 |
-
|
| 440 |
-
| | eval index 42 | train index 1234 |
|
| 441 |
-
|---|---|---|
|
| 442 |
-
| task id | `eval-00042` | `train-01234` |
|
| 443 |
-
| reset image sha256 | `a1447fd8` | `5494a500` |
|
| 444 |
-
| look image sha256 | `1c95af16` | `0850e712` |
|
| 445 |
-
| move distance | 27.778734090271577 m | 26.196924979479498 m |
|
| 446 |
-
| reward, distance | 0.0, 7463.03 km | 0.0, 12647.18 km |
|
| 447 |
-
|
| 448 |
-
The guess map is **pixel-identical** too: mean absolute difference 0.000/255
|
| 449 |
-
over 1020x390. Only the PNG encoding differs, not the content.
|
| 450 |
-
|
| 451 |
-
The one fragile part is the street layer, which comes from Overpass. Overpass
|
| 452 |
-
answers a laptop in ~2 s but intermittently returns **504 Gateway Timeout** to
|
| 453 |
-
datacenter egress, which once left the Space rendering coarse
|
| 454 |
-
Natural-Earth-only maps while a laptop drew the full labelled grid. A single
|
| 455 |
-
retry recovers it, and the environment now *reports* the state rather than
|
| 456 |
-
degrading silently: observation metadata carries `street_detail` as `on`, `off`
|
| 457 |
-
or **`unavailable`**, so a poorer map is visible instead of looking like a
|
| 458 |
-
styling choice. Scoring is unaffected either way -- reward is distance-based.
|
| 459 |
-
`GEOGUESSER_OSM_CACHE` points the street cache at a mounted, pre-warmed
|
| 460 |
-
directory for deployments that cannot reach Overpass at all.
|
| 461 |
-
|
| 462 |
-
## Task API
|
| 463 |
-
|
| 464 |
-
The environment implements the core `TaskProvider` protocol, so the Task API
|
| 465 |
-
routes core already registers become live. Task discovery is metadata only; it
|
| 466 |
-
never starts an episode.
|
| 467 |
-
|
| 468 |
-
```bash
|
| 469 |
-
curl localhost:8000/geoguesser_env/splits
|
| 470 |
-
# [{"name":"train","type":"train","num_tasks":3452,"default":true},
|
| 471 |
-
# {"name":"eval","type":"test","num_tasks":200,"default":false}]
|
| 472 |
-
|
| 473 |
-
curl -X POST localhost:8000/geoguesser_env/num_tasks -d '{"split":"eval"}'
|
| 474 |
-
curl -X POST localhost:8000/geoguesser_env/task -d '{"split":"eval","index":12}'
|
| 475 |
-
```
|
| 476 |
-
|
| 477 |
-
```python
|
| 478 |
-
env.list_splits() # [{"name": "eval", "type": "test", ...}, ...]
|
| 479 |
-
env.num_tasks("eval") # 200
|
| 480 |
-
env.get_task("eval", 12) # metadata, no coordinates and no country
|
| 481 |
-
```
|
| 482 |
-
|
| 483 |
-
**Task specs are deliberately truth-free.** They carry `task_index`, `task_id`,
|
| 484 |
-
`split`, `n_frames`, `provider`, `sequence_id` and `offline_ready` — never
|
| 485 |
-
coordinates and never the country. A spec travels to whatever orchestrates a
|
| 486 |
-
run, and a label sitting in a spec can reach a prompt. The true location is
|
| 487 |
-
revealed in observation metadata after the guess, which is the one place it
|
| 488 |
-
belongs. Per-country eval breakdowns therefore come from finished episodes, not
|
| 489 |
-
from `list_tasks`.
|
| 490 |
-
|
| 491 |
-
Splits are invisible to the agent. `RESERVED_TOOL_NAMES` blocks a `reset` MCP
|
| 492 |
-
tool, so there is no way for a policy to see or choose its own task — the
|
| 493 |
-
"agents cannot reset" invariant.
|
| 494 |
-
|
| 495 |
-
## Collecting evals
|
| 496 |
-
|
| 497 |
-
`scripts/collect_eval.py` runs models against a split and records everything
|
| 498 |
-
about each episode. One JSONL line per episode, ~13 KB.
|
| 499 |
-
|
| 500 |
-
```bash
|
| 501 |
-
export ANTHROPIC_API_KEY=...
|
| 502 |
-
python scripts/collect_eval.py --provider anthropic --model claude-sonnet-5 \
|
| 503 |
-
--split eval --max-turns 12
|
| 504 |
-
|
| 505 |
-
# several endpoints in one run, each at its own concurrency
|
| 506 |
-
cp models.example.json models.json # anthropic / openai / HF router / vLLM
|
| 507 |
-
python scripts/collect_eval.py --models models.json --split eval
|
| 508 |
-
```
|
| 509 |
-
|
| 510 |
-
Output lands in `rollouts/<run_id>/`: `episodes.jsonl` plus a `run.json` with
|
| 511 |
-
per-model aggregates (mean reward, median distance, within-1/25/200/750 km,
|
| 512 |
-
parse rate, forced guesses, tokens, latency). Runs are **resumable** — re-run the
|
| 513 |
-
same `--run-id` and it skips episodes already recorded.
|
| 514 |
-
|
| 515 |
-
Each turn carries the camera state **before and after** the action, which is
|
| 516 |
-
what makes a rollout re-renderable: a pan from 0 to 270 degrees can only be
|
| 517 |
-
animated if both ends are known. It also carries the exact prompt, the raw
|
| 518 |
-
reply, separated reasoning, finish reason, token counts, per-attempt errors and
|
| 519 |
-
latency, and the provider's own response id.
|
| 520 |
-
|
| 521 |
-
**Pixels are not stored.** The environment is deterministic and the panoramas
|
| 522 |
-
are local, so a renderer replays the state trajectory instead. Storing them
|
| 523 |
-
would cost roughly 14 GB for 200 tasks across six models, all re-derivable. What
|
| 524 |
-
*is* stored is a sha256 per observation, so a replay can be *verified* rather
|
| 525 |
-
than assumed:
|
| 526 |
-
|
| 527 |
-
```bash
|
| 528 |
-
python scripts/verify_replay.py ../../rollouts/<run_id>/episodes.jsonl
|
| 529 |
-
# 3 episodes · 19 view turns checked · 19 match · 0 mismatch
|
| 530 |
-
# every view turn reproduces byte-for-byte
|
| 531 |
-
```
|
| 532 |
-
|
| 533 |
-
That check is not decorative: it exits non-zero on drift, because a video built
|
| 534 |
-
from a mismatched trace looks authoritative and shows something the model never
|
| 535 |
-
saw. Guess maps are excluded from it by construction — they depend on the
|
| 536 |
-
Overpass response of the moment. Use `--save-frames` when you want the literal
|
| 537 |
-
bytes anyway.
|
| 538 |
-
|
| 539 |
-
## Rollout video
|
| 540 |
-
|
| 541 |
-
Two steps, split where the work naturally divides. Python owns
|
| 542 |
-
pixels-from-panoramas, because the gnomonic reprojection already lives here and
|
| 543 |
-
is verified byte-exact against the trace. React owns layout, typography and
|
| 544 |
-
transitions, because that is where iterating on them is pleasant.
|
| 545 |
-
|
| 546 |
-
```bash
|
| 547 |
-
# 1. Frames + timeline.json, straight into the Remotion project's public/
|
| 548 |
-
python scripts/render_rollout.py ../../rollouts/<run>/episodes.jsonl --episode 0
|
| 549 |
-
|
| 550 |
-
# 2. Compose
|
| 551 |
-
cd video && npm install
|
| 552 |
-
npx remotion render Rollout out.mp4 --props=public/rollouts/<slug>/timeline.json
|
| 553 |
-
npx remotion studio # iterate on the composition live
|
| 554 |
-
```
|
| 555 |
-
|
| 556 |
-
**The pan is a real pan.** A `look` from 0 to 270 degrees is not a cut between
|
| 557 |
-
two stills: it renders one intermediate gnomonic reprojection per video frame
|
| 558 |
-
along the *shortest angular path*, eased like a camera rather than swept
|
| 559 |
-
linearly — 350 to 10 degrees pans +20, not -340. Zoom interpolates the field of
|
| 560 |
-
view the same way. Verified: a `look` segment produces 29 distinct images, a
|
| 561 |
-
`zoom` 24, with no duplicates.
|
| 562 |
-
|
| 563 |
-
Layout is a large panorama viewport with a HUD of the state the agent is acting
|
| 564 |
-
on (heading, fov, actions left, cost), and a trace pane that reveals turns as
|
| 565 |
-
they happen — the active turn carries the model's raw reply, past turns recede.
|
| 566 |
-
Then a score card with the guess against the truth.
|
| 567 |
-
|
| 568 |
-
Fidelity is checked before anything is composed: every view keyframe is
|
| 569 |
-
re-rendered at the size the model saw and compared to the trace's sha256, and a
|
| 570 |
-
mismatch aborts. A video that looks authoritative while showing something the
|
| 571 |
-
model never saw is worse than no video.
|
| 572 |
-
|
| 573 |
-
## Reproducibility
|
| 574 |
-
|
| 575 |
-
```python
|
| 576 |
-
env.reset(split="eval", index=7) # exact task, byte-identical -> GRPO, eval
|
| 577 |
-
env.reset(split="train", seed=42) # tasks[42 % n_tasks] -> replay
|
| 578 |
-
env.reset() # random task in the default split, split and
|
| 579 |
-
# index both recorded in metadata -> UI
|
| 580 |
-
```
|
| 581 |
-
|
| 582 |
-
`task_index=` still works as an alias for `index=`, so trajectories recorded
|
| 583 |
-
before splits existed still replay. The split is recorded in observation
|
| 584 |
-
metadata: without it a bare index is ambiguous across three indexes, and a
|
| 585 |
-
trajectory stops being replayable.
|
| 586 |
-
|
| 587 |
-
Byte-identical repeats hold because panorama bytes come from a local cache
|
| 588 |
-
rather than an expiring CDN URL, reprojection is pure numpy with integer
|
| 589 |
-
sampling, and the initial heading is pinned to each panorama's own
|
| 590 |
-
`compass_angle`.
|
| 591 |
-
|
| 592 |
-
An eval score is only meaningful alongside its provenance — the env version,
|
| 593 |
-
the task index, and `GEODATA_VERSION` from `server/render/minimap.py`, since
|
| 594 |
-
the bundled vectors determine the reverse-geocode text the agent sees.
|
| 595 |
-
|
| 596 |
-
## Training and collection
|
| 597 |
-
|
| 598 |
-
The environment plugs into `openenv.core.harness`, so a rollout function and a
|
| 599 |
-
collector come for free:
|
| 600 |
-
|
| 601 |
-
```python
|
| 602 |
-
from geoguesser_env import GeoGuesserEnv
|
| 603 |
-
from geoguesser_env.harness import GeoGuesserSessionFactory, load_tasks
|
| 604 |
-
|
| 605 |
-
tasks = load_tasks("tasks/train_pano_v3.jsonl", repeat=16, split="train")
|
| 606 |
-
factory = GeoGuesserSessionFactory(
|
| 607 |
-
lambda: GeoGuesserEnv(base_url="http://localhost:8000")
|
| 608 |
-
)
|
| 609 |
-
```
|
| 610 |
-
|
| 611 |
-
See `examples/geoguesser_rollout.py` for a scripted rollout and
|
| 612 |
-
`examples/geoguesser_collect.py` for JSONL collection with resume.
|
| 613 |
-
|
| 614 |
-
## Human play
|
| 615 |
-
|
| 616 |
-
A five-round game, 5,000 points a round on the same curve the environment
|
| 617 |
-
rewards, so a human score is directly comparable to GeoGuessr intuition and to
|
| 618 |
-
the agent's reward (both are shown).
|
| 619 |
-
|
| 620 |
-
**A round is one episode with one guess.** The five-round game is a UI wrapper
|
| 621 |
-
around five separate episodes; the environment itself never accepts more than
|
| 622 |
-
one guess, because `submit_guess` is terminal.
|
| 623 |
-
|
| 624 |
-
The page plays through the environment rather than simulating it. It opens the
|
| 625 |
-
same WebSocket session API a client uses, calls `reset(task_index=...)`, and
|
| 626 |
-
sends every pin, look, zoom and move as a real charged step — so the step
|
| 627 |
-
counter, the accumulated cost and the final reward are the environment's own
|
| 628 |
-
numbers, not the browser's. A side panel shows the observation stream an agent
|
| 629 |
-
would receive, including the environment's own rendered map and views.
|
| 630 |
-
|
| 631 |
-
Note that plain REST `/step` builds a fresh environment per request, so a
|
| 632 |
-
stateful episode has to run over `/ws`; the Python client does this already.
|
| 633 |
-
|
| 634 |
-
```bash
|
| 635 |
-
uv run --project . server
|
| 636 |
-
# then open http://localhost:8000/geoguesser/play
|
| 637 |
-
```
|
| 638 |
-
|
| 639 |
-
The page stands alone at `/geoguesser/play` and is also embedded in the Gradio
|
| 640 |
-
playground's **Custom** tab when the web interface is enabled:
|
| 641 |
-
|
| 642 |
-
```bash
|
| 643 |
-
ENABLE_WEB_INTERFACE=true uv run --project . server # http://localhost:8000/web/
|
| 644 |
-
```
|
| 645 |
-
|
| 646 |
-
Pick a split and an episode with the `reset(split=)` and `reset(index=)`
|
| 647 |
-
controls above the game and press
|
| 648 |
-
**load episode**, or **random episode** — the same call an eval harness makes,
|
| 649 |
-
so you can replay exactly the episode an agent saw. Those controls live on the
|
| 650 |
-
Gradio side because choosing a task is orchestration, not something the player
|
| 651 |
-
does mid-round; the page itself reads `?task=` from its URL, so
|
| 652 |
-
`/geoguesser/play?task=42` opens that episode directly.
|
| 653 |
-
|
| 654 |
-
Drag to look around and scroll to zoom (free, for orientation). The `look()`
|
| 655 |
-
and `zoom(30)` buttons run charged environment steps and show what the agent
|
| 656 |
-
sees. Arrows, or the arrow keys, walk the road — the main view follows, keeping
|
| 657 |
-
your heading. **M** toggles a larger map, **T** the trace panel, **Enter**
|
| 658 |
-
submits and then advances. On submit the map takes the screen and draws the
|
| 659 |
-
line between guess and truth, exactly like the game; the result bar shows
|
| 660 |
-
distance, points, env reward and the true location, and a scoreboard breaks
|
| 661 |
-
down all five rounds at the end.
|
| 662 |
-
|
| 663 |
-
Panoramas are rendered by Pannellum and the map by MapLibre over OpenFreeMap
|
| 664 |
-
tiles — no API key, no request limits. The imagery credit line names the
|
| 665 |
-
Mapillary contributor, which the CC-BY-SA licence requires.
|
| 666 |
-
|
| 667 |
-
The page has to be a standalone document rather than a Gradio `gr.HTML`
|
| 668 |
-
fragment: `gr.HTML` inserts markup without executing `<script>` tags, so the
|
| 669 |
-
viewers never initialise and the panel renders blank with no error anywhere.
|
| 670 |
-
|
| 671 |
-
Extra routes, all local:
|
| 672 |
-
|
| 673 |
-
| Route | Returns |
|
| 674 |
-
|-------|---------|
|
| 675 |
-
| `/geoguesser/play` | the play page |
|
| 676 |
-
| `/geoguesser/tasks` | `{"n_tasks": N}` |
|
| 677 |
-
| `/geoguesser/task/{i}` | task metadata, including ground truth for the human UI |
|
| 678 |
-
| `/geoguesser/pano/{i}` | the starting equirectangular panorama |
|
| 679 |
-
| `/geoguesser/pano/{i}/{frame}` | one frame's panorama, so the viewer follows `move()` |
|
| 680 |
-
|
| 681 |
-
The human map uses live tiles; the agent's map stays the offline Natural Earth
|
| 682 |
-
render, so the agent keeps a determinism the browser does not need. Note that
|
| 683 |
-
`/geoguesser/task/{i}` exposes ground truth — it exists for a person playing in
|
| 684 |
-
their own browser, and agent observations still withhold it until the guess.
|
| 685 |
-
|
| 686 |
-
## Configuration
|
| 687 |
-
|
| 688 |
-
| Variable | Default | Meaning |
|
| 689 |
-
|----------|---------|---------|
|
| 690 |
-
| `GEOGUESSER_TASKS_EVAL` | `tasks/eval_pano_v3.jsonl` | Frozen eval split |
|
| 691 |
-
| `GEOGUESSER_TASKS_TRAIN` | `tasks/train_pano_v3.jsonl` | Training split |
|
| 692 |
-
| `GEOGUESSER_DEFAULT_SPLIT` | `train` | Split `reset()` uses when none is named |
|
| 693 |
-
| `GEOGUESSER_INDEX` | `tasks/pano_v1.jsonl` | Legacy single index, used only when no split resolves |
|
| 694 |
-
| `GEOGUESSER_CACHE` | `data/panos` | Panorama cache directory |
|
| 695 |
-
| `GEOGUESSER_EPISODE_MODE` | `agentic` | `agentic`, `single_shot` or `nmpz` |
|
| 696 |
-
| `GEOGUESSER_MAX_STEPS` | `24` | Actions before the episode is cut off |
|
| 697 |
-
| `GEOGUESSER_REWARD_MODE` | `coords` | `coords` or `country_only` |
|
| 698 |
-
| `GEOGUESSER_HIERARCHICAL` | `0` | Add country and region partial credit |
|
| 699 |
-
| `GEOGUESSER_VIEW_SIZE` | `640` | Edge length of rendered views |
|
| 700 |
-
| `GEOGUESSER_ALLOW_FETCH` | `1` | Whether a cache miss may reach the API |
|
| 701 |
-
| `GEOGUESSER_HIRES_ZOOM` | `1` | Render views at or below 45 deg fov from the original |
|
| 702 |
-
| `GEOGUESSER_STREET_DETAIL` | `1` | Fetch real OSM streets below 0.35 deg. Governs Overpass only, independent of `ALLOW_FETCH`, and caches to local disk |
|
| 703 |
-
| `GEOGUESSER_REVEAL_MAP` | `1` | Draw the guess-versus-truth map; `0` is 3x faster for training |
|
| 704 |
-
| `GEOGUESSER_OSM_CACHE` | `data/geo/osm_cache` | Street-window cache; point at a pre-warmed mount where Overpass is unreachable |
|
| 705 |
-
| `MAPILLARY_API_KEY` | — | Needed by the builder, and only on a cache miss |
|
| 706 |
-
|
| 707 |
-
## Data
|
| 708 |
-
|
| 709 |
-
Three splits, carved from one 3,673-task pool so contamination is enforced
|
| 710 |
-
exactly once, at split time, rather than reasoned about across two harvests:
|
| 711 |
-
|
| 712 |
-
| Split | Type | Tasks | Countries | Offline |
|
| 713 |
-
|---|---|---|---|---|
|
| 714 |
-
| `eval` | `test` | 200 | 73, capped at 4 each | all 24 frames mirrored |
|
| 715 |
-
| `train` | `train` | 3,452 | 132 | all 24 frames mirrored |
|
| 716 |
-
| `random` | `validation` | 1.2M pool rows | global | no, fetches on demand |
|
| 717 |
-
|
| 718 |
-
Separation follows the OSV-5M rule: no shared `sequence_id`, and no training
|
| 719 |
-
task within 1 km of an eval task. Frames sit ~3.3 m apart, so holding out an
|
| 720 |
-
image while keeping its neighbour holds out nothing. The split script verifies
|
| 721 |
-
its own work and exits non-zero if either rule is violated — the committed
|
| 722 |
-
split reports 0 shared sequences and a closest train task 1.07 km away.
|
| 723 |
-
|
| 724 |
-
| | |
|
| 725 |
-
|---|---|
|
| 726 |
-
| Frames per task | 23.2 mean (8 min, 24 max), ~3.3 m apart |
|
| 727 |
-
| Eval index | 1.2 MB, committed |
|
| 728 |
-
| Train index | 20 MB, in the Storage Bucket |
|
| 729 |
-
| Imagery | 22 GB for 86k frames, 0.26 MB mean per frame |
|
| 730 |
-
|
| 731 |
-
`eval` is committed because a frozen benchmark belongs in version control,
|
| 732 |
-
where a change to it shows up in review. The training index and the imagery
|
| 733 |
-
live in a Storage Bucket, mounted read-only at `/data` on a Space.
|
| 734 |
-
|
| 735 |
-
The `random` split is **not yet implemented** — the plumbing takes arbitrary
|
| 736 |
-
named splits, but the pool-backed sampler is still to come.
|
| 737 |
-
|
| 738 |
-
Each index is self-contained: every frame's coordinates, heading and capture
|
| 739 |
-
date live in the JSONL, so the movement graph resolves offline.
|
| 740 |
-
Only image bytes are fetched, and only on a cache miss, because Mapillary
|
| 741 |
-
`thumb_*_url` values are expiring signed URLs that cannot be stored.
|
| 742 |
-
|
| 743 |
-
Coverage is uneven and worth knowing about. Probing 45 Street-View
|
| 744 |
-
coordinates found any Mapillary imagery at 21 and a 360-degree panorama at
|
| 745 |
-
only 7, heavily clustered. Panorama-first discovery is therefore the only
|
| 746 |
-
approach that works — roughly 5% of probe points yield a usable sequence, so
|
| 747 |
-
reaching 100 tasks took two passes with different seeds, merged by
|
| 748 |
-
`scripts/merge_task_indexes.py`. Africa and Oceania are thin because 360-degree
|
| 749 |
-
contributors are; that is a property of the source, documented rather than
|
| 750 |
-
papered over.
|
| 751 |
-
|
| 752 |
-
## Known gaps versus the real game
|
| 753 |
-
|
| 754 |
-
Movement follows captured sequences and stops where one ends. There is no
|
| 755 |
-
multi-round cumulative score, no wall-clock timer (a step budget stands in for
|
| 756 |
-
it), and no satellite layer on the guess map. Coverage hints and web search are
|
| 757 |
-
deliberately excluded: the first is a crutch, the second turns the task into
|
| 758 |
-
retrieval.
|
| 759 |
-
|
| 760 |
-
See [DESIGN.md](DESIGN.md) for the reasoning behind these choices.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/openenv_geoguesser_env.egg-info/SOURCES.txt
DELETED
|
@@ -1,29 +0,0 @@
|
|
| 1 |
-
README.md
|
| 2 |
-
__init__.py
|
| 3 |
-
client.py
|
| 4 |
-
harness.py
|
| 5 |
-
models.py
|
| 6 |
-
pyproject.toml
|
| 7 |
-
./__init__.py
|
| 8 |
-
./client.py
|
| 9 |
-
./harness.py
|
| 10 |
-
./models.py
|
| 11 |
-
openenv_geoguesser_env.egg-info/PKG-INFO
|
| 12 |
-
openenv_geoguesser_env.egg-info/SOURCES.txt
|
| 13 |
-
openenv_geoguesser_env.egg-info/dependency_links.txt
|
| 14 |
-
openenv_geoguesser_env.egg-info/entry_points.txt
|
| 15 |
-
openenv_geoguesser_env.egg-info/requires.txt
|
| 16 |
-
openenv_geoguesser_env.egg-info/top_level.txt
|
| 17 |
-
server/__init__.py
|
| 18 |
-
server/app.py
|
| 19 |
-
server/geoguesser_environment.py
|
| 20 |
-
server/gradio_ui.py
|
| 21 |
-
server/parser.py
|
| 22 |
-
server/scoring.py
|
| 23 |
-
server/backends/__init__.py
|
| 24 |
-
server/backends/base.py
|
| 25 |
-
server/backends/panorama.py
|
| 26 |
-
server/render/__init__.py
|
| 27 |
-
server/render/minimap.py
|
| 28 |
-
server/render/pano.py
|
| 29 |
-
tests/test_geoguesser_env.py
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/openenv_geoguesser_env.egg-info/dependency_links.txt
DELETED
|
@@ -1 +0,0 @@
|
|
| 1 |
-
|
|
|
|
|
|
geoguesser_env/openenv_geoguesser_env.egg-info/entry_points.txt
DELETED
|
@@ -1,2 +0,0 @@
|
|
| 1 |
-
[console_scripts]
|
| 2 |
-
server = geoguesser_env.server.app:main
|
|
|
|
|
|
|
|
|
geoguesser_env/openenv_geoguesser_env.egg-info/requires.txt
DELETED
|
@@ -1,14 +0,0 @@
|
|
| 1 |
-
openenv>=0.3.1
|
| 2 |
-
fastapi>=0.115.0
|
| 3 |
-
pydantic>=2.0.0
|
| 4 |
-
uvicorn>=0.24.0
|
| 5 |
-
fastmcp>=2.0.0
|
| 6 |
-
pillow>=10.0.0
|
| 7 |
-
numpy>=1.24.0
|
| 8 |
-
matplotlib>=3.7.0
|
| 9 |
-
|
| 10 |
-
[dev]
|
| 11 |
-
pytest>=8.0.0
|
| 12 |
-
|
| 13 |
-
[ui]
|
| 14 |
-
gradio>=4.0.0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
geoguesser_env/openenv_geoguesser_env.egg-info/top_level.txt
DELETED
|
@@ -1 +0,0 @@
|
|
| 1 |
-
geoguesser_env
|
|
|
|
|
|