Files
Zecheng Zhang 8575b9b2f4 fix(github): hydrate the tree lazily, make build_resource sync again
0.0.5 made build_resource a coroutine so GitHubResource could fetch the
repo tree before returning. That fixed a real defect (the two fetches ran
in __init__ over a blocking urlopen and froze the daemon's loop) but paid
for it with the one function every caller who describes a mount as data
comes through: the YAML loader, the daemon's create/load routes, clone,
and every embedder reaching it through mirage.sdk. The haystack
integration, our first outside consumer, broke on it.

Hydrating lazily removes a round trip instead of adding one. Nothing
seeded the index at build time, so the first readdir ran
ensure_live_index and refetched the whole tree, discarding the one the
constructor had just paid for; measured two git/trees calls where one
does. GitHubResource now names the repository and contacts nothing, and
the tree and default branch arrive through ensure_tree and
ensure_default_branch on first use, each behind a lock so concurrent
first reads cost one request. readdir and read already hydrated through
ensure_live_index; only find, du and grep's narrow read accessor.tree
directly and needed wiring.

normalize_resources now refuses a non-resource and names the mount. The
old failure was 'coroutine' object has no attribute 'set_index', raised
two frames away in install_mounts, naming a method the caller never
called and no mount.

TypeScript stays async: its factory type is uniformly
(config) => Promise<Resource> and two of its backends need it. Recorded
as a deliberate divergence rather than mirrored.
2026-08-15 22:39:02 -07:00

86 lines
2.7 KiB
Python

# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# ========= Copyright 2026 @ Strukto.AI All Rights Reserved. =========
import os
from dotenv import load_dotenv
from mirage import Mount, MountBackend, MountMode, Workspace
from mirage.resource.github import GitHubConfig, GitHubResource
load_dotenv(".env.development")
config = GitHubConfig(token=os.environ["GITHUB_TOKEN"])
# Nothing here is async: the mount names the repo and fetches its tree on
# the first read, which is the point of a FUSE mount.
resource = GitHubResource(
config=config,
owner="strukto-ai",
repo="mirage",
ref="main",
)
with Workspace({
"/github/":
Mount(resource, mode=MountMode.READ, backend=MountBackend.FUSE)
}) as ws:
mp = ws.fuse_mountpoint
print(f"=== FUSE MODE: mounted at {mp} ===\n")
print("--- os.listdir() root ---")
entries = os.listdir(mp)
for e in entries[:10]:
print(f" {e}")
if len(entries) > 10:
print(f" ... ({len(entries)} total)")
print("\n--- os.listdir() python/mirage/ ---")
core = os.listdir(f"{mp}/python/mirage")
for c in core[:10]:
print(f" {c}")
print("\n--- os.listdir() python/mirage/core/ ---")
core_dirs = os.listdir(f"{mp}/python/mirage/core")
for d in core_dirs[:10]:
print(f" {d}")
if len(core_dirs) > 10:
print(f" ... ({len(core_dirs)} total)")
print("\n--- open() + read python/pyproject.toml (first 5 lines) ---")
with open(f"{mp}/python/pyproject.toml") as f:
for i, line in enumerate(f):
if i >= 5:
break
print(f" {line.rstrip()}")
print("\n--- open() + read python/mirage/types.py (first 5 lines) ---")
with open(f"{mp}/python/mirage/types.py") as f:
for i, line in enumerate(f):
if i >= 5:
break
print(f" {line.rstrip()}")
print(f"\n>>> FUSE mounted at: {mp}")
print(">>> Open another terminal and run:")
print(f">>> ls {mp}/")
print(f">>> cat {mp}/README.md")
print(">>> Press Enter to unmount and exit...")
input()
records = ws.ops.records
total = sum(r.bytes for r in records)
print(f"\nStats: {len(records)} ops, {total} bytes")