bbc6ba362f
The current meta-schedule uses a PostProc `RewriteUnboundBlock` to auto-bind blocks to threads. However, it's a post proc, which means there are no search opportunities, and always splits with `factor=1024`. This PR adds a new search rule called `AutoBind` to do a similar thing to bind threads with sampled factors. Also with a corresponding mutator. After applying this rule, we get some positive perf results (on RTX-3080): Element-wise: from 2.76 us to 2.48 us Conv2d Winograd: from 29.45 us to 18.96 us (ansor 22.00 us) Resnet18: from 0.591 ms to 0.531 ms (ansor 0.565 ms)