Commit Graph

2475 Commits

Author SHA1 Message Date
Zilin Zhu 39fa759906 fix docs of threefry_split and threefry_generate (#8035) 2021-05-13 09:22:07 -04:00
Yuchen Jin 43c2ea72bc Rename gpu to cuda, and bump dlpack to v0.5 (#8032) 2021-05-13 09:11:40 -04:00
Matthew Brookhart ed283b82bb support concat in recast (#8028) 2021-05-12 23:43:36 -07:00
anwang2009 0f41d47bb4 Remove minimum seed constraint on XGB Tuner (#7992)
* remove minimum seed

* reset 3rdparty dep

* add items to 'visited', parametrize min seed records

* add comment

* fix lint

* add tests
2021-05-11 10:44:01 -07:00
Josh Fromm c30f099b9e Fix bug with non-fp32 gemm in onnx frontend. (#8011) 2021-05-11 09:16:54 -07:00
alter-xp 2077ee2f9a add onnx reverse sequence op (#7771)
Co-authored-by: xp224797 <xp224797@alibaba-inc.com>
2021-05-10 09:58:26 -06:00
Trevor Morris 4c1a6aa956 [BYOC][TensorRT] Add nn.batch_matmul, nn.layer_norm, erf (#8005) 2021-05-08 09:35:53 -04:00
Manupa Karunaratne c0690496af Improved MLF to contain workspace info (#7938)
* Improved MLF to contain workspace info

Added functionality to calculate workspace, io and constant
memory required by each primfunc and main function. Moreover,
the workspace information required by each primfunc and main
is reported in metadata.json in the Model Library Format(MLF).
- added functionality to record tir and relay primfuncs
- added tests for model_library_format changes

Change-Id: Ib4a8b787345aa35f8a1645e8a648fad84de37bce

* Improved MLF to contain workspace info

* disable AoT for now
* addressing comments

Change-Id: I5f041ec461b02dac6ea9c96ea50eb400d55eef53

* Improved MLF to contain workspace info

* addressed comments
* added aot executor support

Change-Id: I9b54a7939d8ccb3c6ce0454f0fe62866ac66eb5c

* Improved MLF to contain workspace info

* removed redundant utils.py

Change-Id: I256dd88fab31a595bf9509bd1c4ab59b0c145b1e

* Improved MLF to contain workspace info

* removed redundant ffi api

Change-Id: I9ad6795aa839edfdfd05b902d4531fb0a20e894d
2021-05-07 14:54:40 -07:00
Siyuan Feng 4122a6aed4 [TensorIR] CreatePrimFunc from TE (#7987)
Co-authored-by: Tianqi Chen <tqchen@users.noreply.github.com>
Co-authored-by: Wuwei Lin <wuwei@apache.org>
Co-authored-by: Ruihang Lai <lairuihangdongdong@qq.com>
2021-05-07 13:31:31 -07:00
zackcquic 254563a314 [RELAY] Enable registering op with python (#8002)
Add a new API register_op

Note: Implementing a op by pure python is still limited:
  1. Custom type relation (add_type_rel()) is still not
     available in python.

  2. Setting number inputs (set_num_inputs()) needs
     plevel > 128 in python.
     (see tests/python/relay/test_ir_op.py)
2021-05-07 11:29:02 -04:00
Yuchen Jin 8d9a1dfe77 [DLPACK] Support the new python array api with DLPack (#7993)
* [DLPACK] Support the new python array api with dlpack

* Fix lint
2021-05-06 22:51:49 -07:00
Trevor Morris d18186757c [Frontend][Keras] Support nested layers recursively in keras frontend (#7949)
* Support nested layers recursively in keras frontend

* Fix lint

* Fix issue

* Fix formatting

* Fix unit test
2021-05-06 11:13:24 -06:00
Tristan Konolige bbf7fdb8c0 [FIX] Fix autoscheduler tuning on sparse matrices where there are multiple with the same shape (#7974)
* [FIX] Fix autoscheduler tuning on sparse matrices where there are multiple with the same shape

* formatting

* remove unreachable code
2021-05-06 10:58:48 +08:00
Giuseppe Rossini f85cab2052 [AOT] Introducing AOT in TVM (#7785)
* [AOT] Introducing AOT in TVM

This change adds the code generation and minimal runtime API to use the
Ahead Of Time (AOT) compilation flow. The main logic is contained in:

- src/relay/backend/aot_codegen.cc

Which produces a TIR PrimFunc traversing the Relay graph

The runtime interface (authored by @mousius) leaves a gap for future
iterations using platform-specific features from RTOS.

Currently AOT runs successfully on x86 in a host OS, running these
tests on micro is coming soon.

This PR is based on the RFC described here: https://discuss.tvm.apache.org/t/implementing-aot-in-tvm/9206

Co-authored-by: Christopher Sidebottom <Christopher.Sidebottom@arm.com>
Change-Id: I9f731c953231f129e1472298915dddc01788efd7

* Rebasing 2

Change-Id: Ia0a533a49960f1cb4bf3c3833511e539cf7c459f

* Applying comments/refactoring

Change-Id: Iea1832355f8b1d4c921d02c6b4ceec7db3a681c1

* Fixing comments + refactoring - 2

Change-Id: I7200cc17b297e42bf67dcdef6f643e86991ca0a8

* fix linting

Change-Id: Iba6544ac7101595696b352b8702345cf916625f6

* fix linting - 2

Change-Id: I7f80d16005f2c621d37a9aae2cbbd61df0277cbe

* fix linting - 3

Change-Id: I7a1ba40afeea46d5f122563a20cd4b2f08751a1e

* fix tests

Change-Id: I1297ccc54dd6d93647f421e0beb226f410bf73f5

* Addressing comments - 3

Change-Id: Id25d1382c30d6d0a0013b5e8986fb8cd886666dc

* Addressing comments - 4

Change-Id: Ibe29676abe3b75161b5a0903e007118a8318d862

* fix tests - 2

Change-Id: I2117f9d4392bfd87102ecbef0993c8b320f479a0

* fix tests - 3

Change-Id: Ic0373543b0f9a54dbd4dc32d428272f7293200ba

* fix tests - 4

Change-Id: I8a6f229c9a3a9e169779c8d49cbfa3f473348b1f

* Addressing comments - 5

Change-Id: Ib9ccd07c87392034a21b2eb70955d0b091b780f1

* fix tests - 5

Change-Id: I4b13c3b548ced414991e83072e9e6fc99b64f939

* fix tests - 6

Change-Id: Id5af1f778ae25bc60849cc054a605181c1b7a765

* addressing comments - 6

Change-Id: Id94a2bbcaae891f9498d41be538f13a952f55b81

* fix linting - 4

Change-Id: I371a0aa5b81824b5a3a1278fac22ace57832027a

* add missing file

Change-Id: If359bef96dd0773ead4f75f0d9f5234276347e2d

* fix build

Change-Id: I73fc1feb6f7b5d454a528e3289228484dc2b07d5

* addressing comments - 7

Change-Id: I7f908f3908ffc77e408391f62edcc06f2600c6c2

* addressing comments - 8

Change-Id: I90bced4e18259a6d42e6a406d93958e204f3859e

* rebasing

Change-Id: Id28751b069bd046f00faee301b2b446b2ea4fab8

* Addressing comments - 9

Change-Id: I06c9f280de0a9bf0ca5545bbbbfcc70cb66831b3

* fix tests - 7

Change-Id: I739f29779862f05def36e5f3e0722019596d17f8

* Addressing comments - 9

Change-Id: Ie736f40a5225f4e56e79006753d7732127da5408

* Applying comments + fixing tests

Change-Id: I83e16068b93aaccc7a86b79d42f13328bc76b53d

* Applying comments - 10

Change-Id: I443d72f53913849f3c28fd6e416162d1ca99e647

* Addressing comments - 11

Change-Id: I7fefbd0076949b9c38d0abbf2759ebf1502de330

* Addressing comments - 11

Change-Id: Iad028144d7b394b2dd2fce41a35ca689d1680200

* fix tests - 7

Change-Id: I14286e665dcdba1e9bc10bb5a27dd6ced50372b0

* fixing tests -8

Change-Id: I7b4c966da9680870ceda1704c749ee3bdc751926

* fixing tests - 9

Change-Id: Icf62128a604998ed1b7d5af4cbeadf7d39196d0b

Co-authored-by: Christopher Sidebottom <Christopher.Sidebottom@arm.com>
2021-05-05 11:21:22 -07:00
Xingyu Zhou ae31a3399b [Frontend][Tensorflow]add batch_dim support for gatherV2 (#7951)
* add batch_dim support

* fix lint

* add check for num of arguments for topi.take

* fix gpu test cases

* add check for batch_dims in take_grad
2021-05-05 09:37:00 -07:00
Trevor Morris 26a5e299be [Frontend][Keras] Fix Dense with 3d inputs (#7753)
* Fix keras rnn dense

* Fix unit test

* Fix unit test
2021-05-04 18:46:26 -06:00
Tristan Konolige 9070c65889 [SPARSE] Improve sparse performance on ROCM (#7935)
* [SPARSE] Improve sparse performance on ROCM

The current sparse dense gpu kernel uses warp level storage to handling
caching of data. Warp level storage uses shuffle intrinsics, which are
slow on rocm (because they actually read and write to shared memory).
Rocm does provide intrinsics to do the correct memory management, but
they are not available through tvm. Instead this PR switches to using
shared memory on rocm devices. Performance is about 2x faster.

* default to shared mem

* formatting

* formatting
2021-05-05 04:25:06 +09:00
AndrewZhaoLuo 38e0bbed7d [ONNX][TOPI][Relay]Support dilations in pooling operators (#7928)
* change more pooling operators

dilations -> dilation to match old field names in conv

fix python interface into new relay nodes

fix order of arguments

update type relation for dilations

change topi interface to use dilations

* spooky, there are two implementations! Change to 1 topi

use generic poolnd instead of 2d implementation for topi

remove old pooling topi

* rename pool --> pool2d in topi

change pool -> pool2d, make topi tests work now

make op level 2 pass with interface changes

fix dilation being hardcoded to 1

proper calculation for avgs among dilations

proper avg pool padding behavior

change name of pool test to pool2d test

* add poolnd baseline implementation

more fixes to edge cases for poolnd, delete old versions

replace topi tests with new baseline python version

clean up tests

make tests more readable kind of

add dilation topi tests FINALLY

remove see_pool.py

remove dilation from grad

* fix subtle implementation detail between topi and baseline python pool op

* rewrite tests to be more generic for relay pooling ops

add relay dilation tests, FINALLY

add some comments to testing code

linting and formatting

add ASF header

make 10/10 for black formatting lol

more appeasing the formatting gods

wow

add parameters to documentation

fix test import

Jostle CI

fix more broken unit tests using old version of pool

fix wrong var used for bound calc

add dilation to arm tests

add docstring to python make funcs

* fix pattern utils out of place args

* properly forward more tests to use dilations in pooling

formatting

more formatting

relax constraints on test to make it pass

relax more constraints

fix some pytorch frontend errors

fix error

better test conditions

jostle build

* fix padding bug with ceil mode

jostle build

cleaner pool condition

remove see_pool.py again

* add dilations field to onnx importer

blacking files

black file

* address matthew's comments
2021-05-04 10:24:21 -06:00
Josh Fromm 18ce8e4b82 [TVMC] A simplified TVMC API for python scripting (Part 1). (#7823)
* Introduce new TVMC Python API.

* Add simple testing model.

* Split result utils into stand-alone file.
2021-05-04 07:59:34 -07:00
Trevor Morris 0e3d850983 [BYOC][TensorRT] Fixes for explicit batch mode, Support reduce to scalar, Support split op (#7967) 2021-05-04 01:06:32 -07:00
Tianqi Chen 284faf241f [RPC] Make tracker jupyter friendly (#7961)
This PR uses the PopenWorker to handle the tracker start up
and makes the tracker jupyter friendly.
2021-05-03 13:41:02 -07:00
Siyuan Feng 22c8f8cca5 [TensorIR][Pass][M1c] FlattenBuffer (#7962)
Co-authored-by: Tianqi Chen <tqchen@users.noreply.github.com>
Co-authored-by: Ruihang Lai <lairuihangdongdong@qq.com>
2021-05-03 10:28:28 -07:00
Mehrdad Hessar 8c56ce3b90 [Graph Executor Debugger] Fix parameter dump (#7903)
* remove debug mode

* reformat

* format

* address comments

* add single call for all layers

* fix test

* revert

* address comments

* address comments

* fix rerun node

* fix error

* format

* raise error on array()

* fix java

* Revert "fix java"

This reverts commit c4cf952dbc5c9c32d65ef0ca05d6ecbb5c06d5aa.

* bring back for java api

* fix error

* cleanup

* format

* rm redundancy

* add last execution track

* trigger build

* address comments

* format

* fix name overlap

* Trigger Build

* trigger build

* trigger

* trigger
2021-05-03 10:17:05 -07:00
Josh Fromm cea7cf16ce Improve dtype detection in loop to fix onnx tests. (#7934) 2021-05-03 09:16:27 -06:00
Matthew Brookhart c380a699e2 [ONNX] More Unit Tests! (#7956)
* support same lower and maxpool in autopad

* fix isinf tests

* lower tolerance on roialign test becuase the onnx result is cropped to 4 decimal places

* slow support for bottom-k

* throw with nullptr in gathernd and scatternd, fix typo

* fix lint

* fix a copy typo
2021-05-03 13:02:29 +09:00
Tianqi Chen dd5379f88a [NVCC] Bugfix nvcc command tool that relies on the compile time env (#7964)
* [NVCC] Bugfix nvcc command tool that relies on the compile time env

* Update python/tvm/contrib/nvcc.py

Co-authored-by: Cody Yu <comaniac0422@gmail.com>

Co-authored-by: Cody Yu <comaniac0422@gmail.com>
2021-05-03 13:01:15 +09:00
Robert Kimball f4a680d80c Replace 0.0.0.0 with 127.0.0.1 for client connections (#7766)
* Rename references to 0.0.0.0 to localhost. Also change references to 127.0.0.1 to localhost so that all references are consistent. 0.0.0.0 is not the same as localhost.
2021-05-02 07:43:12 -04:00
Andrew Liu 08d73454fd fix Relay build docstring (#7963) 2021-05-01 22:31:56 -07:00
Josh Fromm 1ec86605a0 [Topi] Fix arm_cpu bitserial schedule with elemwise ops. (#7929) 2021-05-01 17:28:40 -04:00
Dmitriy Smirnov 63e9e5ad0b [RPC] Bugfix. Removed server forcing IPv4 protocol (#7953)
Removed forcing IPv4 protocol from python RPC server implementation
to be in correspondence with the RPC client implementation which is
used `platform default`. This had led to situation when "localhost"
was translated as 127.0.0.1 for the server (IPv4 protocol was used),
but the client translated it as "::1" and was trying to connect to
server using IPv6 protocol and was getting "ECONNREFUSED 111 Connection refused".

Change-Id: I44802eb1ea78f3b36ac664f0be7237e62084c234
2021-05-01 17:24:33 -04:00
Christoph Gerum 6d555b6b43 Correctly build with -runtime=c without -system-lib (#7954) 2021-05-01 07:44:02 -04:00
Andrew Liu dc1f189207 [AutoTVM] [TOPI] Support AutoTVM for int4 tensorcore (#7831)
* initial

* int4 asnumpy

* remove

* random test

* format

* random

* remove unused import

* change dist range

* add fuse_pack in

* random engine

* reformat

* remove import

* add cuda context

* refactor code
2021-05-01 16:27:36 +08:00
Dmitriy Smirnov bf20107ffe [Tophub] Race condition fixed in folder creation (#7940)
* [Tophub] Race condition fixed in folder creation

Tophub download routines switched to Pathlib's `Path.mkdir`
in order to avoid race conditions in creation of folders
2021-04-30 07:52:25 -04:00
Tianqi Chen 62309e51f5 [TIR][TRANSFORM] Return value support in tir.tvm_call_packed (#7932)
This PR fixes the return value support in tir.tvm_call_packed

- Clarified the semantics of the intrinsics
- Fix a problem when lowering call packed with nested scopes(let bindings)
- Added regression tests to cover the changes
2021-04-30 10:04:58 +08:00
Matthew Brookhart 8fce89500c [TOPI][RELAY][ONNX] Scatter ND (#7927)
* passing topi tests

* passing relay tests, needs better shape checking still

* support ONNX operator

* add shape checking back in

* fix lint

* update docstring
2021-04-28 13:13:06 +08:00
Siyuan Feng dee3133c54 [TensorIR][PASS] CompactBufferAllocation (#7923)
Co-authored-by: Tianqi Chen <tqchen@users.noreply.github.com>
Co-authored-by: Junru Shao <junrushao1994@gmail.com>
Co-authored-by: Cody Yu <comaniac0422@gmail.com>
2021-04-28 13:11:43 +08:00
Mike He f681359b2e [FIX] skip_conv_layers will affect quantization of nn.dense (#7795)
* [FIX] `skip_conv_layers` will affect quantization of `nn.dense`

* [ add ] quantization test case for dense & conv2d

* [ fix ] reformat

* [ reformat ] test file
2021-04-27 00:29:04 -04:00
Y 82fecbfa66 [CodeGenC] Fix bugs when calling extern functions (#7911) 2021-04-26 08:28:35 -04:00
PENGUINLIONG 54fdcc52ae Enable StackVM in AutoTVM (#7897) 2021-04-26 08:21:58 -04:00
Duke Wang f92d7fc629 [BugFix]: Convert tuple to int (#7880)
* DEBUG: Convert tuple to int

* Add test cases for test_conv2d_hwcn()

* CI pass
2021-04-25 15:21:37 +08:00
masahi a741652f37 [TIR][SPIR-V] Fix computing clz on int64 input for vulkan (#7913)
* Fix computing clz on int64 input for vulkan

* rebase fix

Co-authored-by: masa <masa@pop-os.localdomain>
2021-04-24 12:36:22 -07:00
srinidhigoud fad10d7914 [Frontend][Tensorflow] SelectV2 and BroadcastArgs op support for tf2 models (#7901) 2021-04-24 10:11:31 -07:00
Xiyou Zhou 251f57160a [Target][Lowering] Update Op Intrinsic Lowering Mechanism And Intrinsic Lowering Pass (#7809)
This PR updated the intrinsic lowering pass to support the new op registry and avoid overloading the global tvm registry. Meanwhile, it kept the fallback mechanism to find the most suitable lower intrinsic function, e.g., llvm.FLowerIntrinsic vs. default.FLowerIntrinsic. All previous op registration are ported to new functions, and some missing ops would be added in separate PR.
2021-04-23 23:02:38 -07:00
Matthew Brookhart de0bff81fb [ONNX] Support importing Conv with missing attributes (#7899)
* [ONNX] Support importing Conv with missing attributes

* fix removal of attributes ONLY when they are default and for autopad

* move comment to the right place
2021-04-23 10:04:28 -07:00
masahi 373bac25b0 [Relay] Shape func fix for all_class_nms and where op (#7910)
* fix missing cast to int64 in all_class_nms shape func

* fix scalar in where shape func

* add add test

* update test

* minor fix

* add where scalar shape func test
2021-04-23 11:03:59 -06:00
Matthew Brookhart 0e2d5ea479 [ONNX] Support NMS Center Box (#7900)
* [ONNX] Support NMS Center Box

* fix silly mistake in contional
2021-04-22 09:45:03 -06:00
Josh Fromm 8a74388258 [Relay][ONNX] 1-D global and adaptive pooling. (#7906)
* 1D adaptive pooling added and tested.

* Apply formatting.

* Add onnx integration and tests.

* Busted by lint.
2021-04-22 09:44:43 -06:00
Lunderberg 46e0634fb7 [Runtime] Driver version + consistent clock speed units (#7867)
* Added kDriverVersion to DeviceAttrKind, implemented for VulkanDeviceAPI.

The vulkan backend has had inconsistencies that look correlated to
drivers used.  This will help in collecting information for
troubleshooting.

* Changed units for OpenCL's clock rate from MHz to kHz, to match Cuda/ROCm.

* [Docs][Runtime] Additional documentation for tvm.runtime.Device, DeviceAPI feature matching

Primarily documentation, with some changes to the OpenCL DeviceAPI to
match available features in cuda/vulkan.

* Added CL_TARGET_OPENCL_VERSION definition, for use with unified OpenCL headers.

Co-authored-by: Eric Lunderberg <elunderberg@octoml.ai>
2021-04-22 12:53:25 +09:00
Matthew Brookhart 1c71a064f2 [ONNX][TOPI][RELAY] Resize refactor (#7883)
* adds rounding mode for nearest neighbor, passing onnx unit tests for nearest neighbor

* passing all linear test. passing all nearest tests except crop and resize, which needs a dynamic implementation of crop and resize

* most of the bicubic tests are working

* fix exclude outside

* remove dead code

* fix lint

* fix defaults to match old implementation

* fix lint

* fix gpu tests

* fix lint again

* change order of operations to prevent GPU rounding errors
2021-04-21 14:35:27 -07:00
Trevor Morris f28c75f45f [Frontend][Tensorflow] Support SAME padding for dynamic h, w when stride == 1 (#7885)
* Support SAME padding for dynamic workloads when stride == 1

* Fix lint

* Fix lint
2021-04-21 11:30:59 -07:00