* add half and mix precision support to cublas backend
* add TensorCore support in CuDNN
* enhance CuDNN support
* address comments and fix lint
* fix
* add fp16 test
* [VM] add a few more API to vm
* [VM][Fix] fix vm convert args
* [VM] a few fixes
* rename fields
* update
* update vm profiler
* x
* add doc
* lint
* fix test
* address comments
Previously, we would rely on the later phases to error out
(often for using too much shared memory). This enables the
checks on the IR that already exist for CUDA and OpenCL also
for ROCm.
* [Relay][Frontend][Tensorflow]Add conv2d_transpose
* add transformation from NHWC to NCHW to compatible with TVM conv2d_transpose implementation
* remove 'dilations' paramater to compitable with TF1.3
* Add qnn conv2d attributes for input_tensor_scale and
kernel_tensor_scale.
The lowering in the tflite frontend loses the input_tensor_scale
and the kernel_tensor_scale by multiplying it and putting it into
the Requantize operation. This means that any graph partitioning
passes or other passes that need to access this information no longer
have it available in the qnn dialect.
regards
Ramana
* Store input tensor scale and Weight tensor scale for Dense as well
As for conv2d, the tflite frontend drops the input tensor
scale and the weight tensor scale from the relay op. Store
it as separate fields in there.
* Fix unintentional tab
* Rename input_tensor_scale to input_scale and kernel_tensor_scale
to kernel_scale for conv2d.
* input_tensor_scale -> input_scale weight_tensor_scale->weight_scale
* Rework dense testcase
And use input_scale and kernel_scale
* Be consistent in use of input_scale and kernel_scale values
* Fixup qnn conv2d tests for input_scale and kernel_scale
* Make pydoc identical between conv2d and dense for weight_tensor
* Fix up conv2d parameters to be in the same order between C++ and python
* Fix ordering of parameters for dense.
* Add input_scale and output_scale to try and satisfy ci gods
* Delete input_scale and kernel_scale.
nn.conv2d does not contain input_scale and kernel_scale. We need
to delete it when lowering it to nn.conv2d.
* Add input_scale and kernel_scale for qnn.conv2d
* AutoTVM: selecting tuning templates when extracting task
Make the procedure of trying new templates easier.
Test: tests/python/relay/test_autotvm_task_extraction.py
* Use dict to match key for topi ops
* fix lint issue
* be more pythonic :)
* Fix constructor pretty printing
* Make Module::HasDef name consistent with API
* Add VM constructor compilation via eta expansion
* Lint
* Fix CI
* Fix failing test
* Address comment
* Retrigger CI
* Retrigger CI
* WIP Run the TF tutorial on TF2
* Remove debugger statement.
* Complete the support for TF2.0's `resize`.
TF2.0 adds a `half_pixel_centers` attribute to the `resize` function in
the image API. This commit completes the hooks in Relay's TF frontend.
At the point of this commit, no new test yet. Also, this commit
addresses solely the `resize` change. Other commits address other
changes in TF2.0.
* Support TF2.0 in the tutorial by using the compat API.
This looks cleaner than trying to detect the TF version.
* Use the TF compat API, so as to support TF2.0.
This is a direct change, relying on the compat API provided by the TF
team.
This code will last as long as the compat API exists, so a
"proper" support for TF1.x and 2.x will require more work in some
future.
* Partial support for EXPLICIT padding introduced in TF2.0.
Explicit padding is a special case in TF2.0 (see reference linked
below). Some models are serialized with that mode, and break TF support
in TVM.
Support is *partial* as EXPLICIT falls back to set padding on the
Relay op, which only supports 2 values. At some point, padding may need
to be extended to support 4 values, but that is out of scope of this
support commit.
Reference on EXPLICIT padding: https://github.com/tensorflow/tensorflow/commit/ec81825aaf7e848d9f8ddffdf1e0d20aebe9172c#diff-1d1c0bb0a880f85b6164f71dbb2f446e
* Guard on checking for optional TF2.0 attribute.
* Do not expect Relay to implement TF-specific attributes.
The `half_pixel_centers` attribute is a new feature in TF2.0. Earlier
commits of mine mistakenly introduce them in the Relay API. This is
probably not what Relay is expected to support, and the semantics of
`half_pixel_centers` is unclear (to me, at least) at this point.
* Remove unclear comment.
CR https://github.com/dmlc/tvm/pull/4104#discussion_r338705742
Addresses #4104
* Changes after review.
Complying without understanding the rationale for now.
* Fix the arguments set mistakenly.
An argument ignored for the wrong operation.
Previously runtime::Module was supported using shared_ptr.
This PR refactors the codebase to use the Object protocol.
It will open doors to allow easier interpolation between
Object containers and module in the future.
* Add Auto TensorCore TensorCore Unit Test
* Rebase to tvm master branch & Add auto tensor core
* Code Refine
* Add tensor core switch by pragma
* Add pragma in tensor core example code
* Get real tile size to replace hard coded 16
* support more than 2 dimensions (e.g. batchmatmul) for buffer bind scope
* support batch matmul
* Move cuda env check to tensor_core.cc
* Coderefine for tensor_core.cc
* Refine comments
* Some refinements of code and comment
* Update TensorCore UT to pass the CPU test
* remove redundant code
* matmul's storage align for different layout
* Add support for differenct position of type cast
* Add formal tutorial for auto tensorcore codegen
* move tensorcore check up to tutorial code
* code and doc refine
* comment out tune_and_evaluate in tutorial
* fix cpplint error
* Batch matmul tuning running but with errors.
* Default x86 schedule as good as before.
* Code Cleanup
* Remove unused argument.
* improved template documentation.
* Silly lint fix
* Removed leftover comment.
* Moved cfg declaration to schedule for batch_matmul
* Moved x86 dense cfg declaration to schedule.
* lint fix
* Removed duplicate cfg declaration in dense.
* Reverted changes to dense.