bbfdab79d9
## Summary - Keep the Python test launcher close to plain `pytest -n auto`, move nightly tests under `tests/nightly/python`, remove obsolete launchers and collection bookkeeping, and partition CPU/GPU jobs with explicit `gpu` marker expressions. - Repair exact-pointer regressions at their owning boundaries: packed raw-string ABI values, CUDA/Metal matrix intrinsic pointers, internal TE extern offsets, MetaSchedule scalar annotations, localized auto-tensorization scope matching, and typed DLTensor fixture fields. - Preserve typed workspace calls in TIR and cast pointer-returning external calls in CodeGenC, covered by a plain-TIRx 1024-byte global workspace that is compiled as C++. - Finish phasing out value-bearing Relax `R.Prim` annotations by requiring an explicit dtype, removing obsolete value-based contracts, and expressing the DISCO rank-dependent slices as explicit scalar `call_tir` inputs. - Gate the distributed callback on the optional DISCO runtime, NCCL, and at least two GPUs so capability-limited jobs skip instead of failing. - Remove the non-demonstrating pointer probe, use direct TVMScript comparison for packed strings, and remove the four designated legacy testing modules. The seven repaired CPU categories cover packed raw strings (7 failures), CUDA/Metal matrix access-pointer types (7), internal TE extern offsets (1), a typed DLTensor fixture (1), MetaSchedule scalar annotations (1), CodeGenC workspace return casts (12), and localized auto-tensorization storage-scope matching (19). ## Validation - Base: `ded6ad8dd212869c881efb5590f8a33fc972728e` - Head: `a7277e86dbcfe0638c8c252d36760859c4ab4297` - All 35 locally available original failing node IDs pass across the focused runs. - The full focused TE, TIR builtin-lowering, and CodeGenC files pass: 61 tests. - The complete touched Relax/TVMScript set plus PlanAndUpdateBufferAllocationLocation passes with 784 passed, 20 skipped, and 1 expected failure. - The DISCO callback collects and skips when its runtime or two-GPU environment is unavailable. - Six direct mapping tests, twelve tensor-core sketches, and the dp4a sketch pass unchanged. - The compiler rebuild, branch-wide pre-commit hooks, and full-range whitespace checks pass. - The 13 broad CBLAS/TFLite nodes remain dependency-gated; their owning TE and generated-C regressions compile. No merge is included in this change.