122be3fb18
* add group_conv2d_transpose_nchw to CUDA backend * simplify significantly, just add groups argument to conv2d_transpose_nchw