Commit Graph

19 Commits

Author SHA1 Message Date
Adam Gibson e1bda9dbfb Maven version updates for lombok for java 25 support (#10243)
* Add files for pr_1_3_maven_poms

* Update Maven POMs for CUDA backend and platform tests

* Remove engineers-guide-to-llms and pp-step-clone directories

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
2025-12-16 18:17:31 +09:00
Adam Gibson 0f94605d5f ADR/publishing prep for renamespacing and release updates (#10212)
* ADR for renamespacing and release updates

* ADR for renamespacing and release updates

* use new secrets for publishing

* use new secrets for publishing

* remove old distribution management

* remove try catch for cuda build

* update string functionality for cuda

* update string functionality for cuda

* update string functionality for cuda

* update string functionality for cuda

* update string functionality for cuda

* update string functionality for cuda

* remove old distribution management

* remove more ossrh references

* update environment variables

* update lombok version

* clean up test dependencies errant linker flags

* get rid of dup flags
2025-05-27 12:02:19 +09:00
Adam Gibson 722592d8a9 Update for publishing to snapshots (#10211)
* fix extra quote inserted by intellij

* Fix path

* fix executable permissions usage

* improve directory resolution and restore some files from the previous PR

* remove old unused dep versions
Update to new march 2024 central release plugin

* update central publishing version

* temp disable mirror

* fix hallucinated version

* update maven invoker version for jetspeed

* remove old build properties
add better directory resolution for copy flatc

* fix helpers sources for cpu (api updates)
perform maven antrun upgrade

* more fixes for onednn helpers

* update more antrun usage

* fix pooling param usage

* remove findbugs on the 2 files it was used on

* debug odd build path issues

* try again

* remove bad ls

* try again

* another attempt

* Add troubleshooting logs again

* Remove specific gcc versions

* update build

* cd to original directory instead

* Remove enforcer check

* remove debug steps

* remove extra steps

* update linux, android and cuda versions

* update arm usage defaults

* remove 32 bit builds

* alternative path

* alternative path get rid of old tools

* get rid of over simplified bootstrap libnd4j download usage

* remove extra cmake

* remove manual path setting

* arm compilation fix

* debug armcompute

* update android-x86_64 openblas

* upgrade gcc

* minor armcompute const changes

* update openblas version

* fix armcompute paths

* remove ls

* update cuda versions and envrionment variables

* fix syntax errors

* fix syntax errors

* fix syntax errors

* fix syntax errors

* another rev

* update msys command

* rev

* fix classifier

* rev android-x86_64

* rev android-x86_64

* update cuda env passing

* fix linux-arm64 syntax error

* cuda rev

* cuda rev

* update arm else if branches for matching

* remove external PS script?

* remove external PS script?

* remove external PS script?

* improve architecture detection

* update list operation conv2d armcompute

* add flatbuffer generated code

* remove old maven auth

* remove old legacny average and accumulate

* remove cache due to 422 error

* ensure armcompute is optional

* ensure armcompute is optional

* disk space clean up on all workflows

* deal with windows service issue

* ensure nvcc is on path

* Add arm64 protoc
Update cuda paths

* convert to use numeric types only mitigating lld errors found by clang

* Add arm64 protoc
Update cuda paths

* fix syntax error

* fix overriding properties causing libnd4j not to be built with cuda

* update windows version

* try to update cl.exe paths

* update debugging for nvcc

* update cuda flags to work with linux/windows

* remove extra cuda install

* fix duplicate sources

* more unsupported compiler changes

* remove unneeded source and javadoc

* better cudnn detection

* update where we put unsupported compiler

* Collapse cmake logic

* Collapse cmake logic

* clean up consolidated file

* flatbuffers fix

* fix hallucinated paths

* generate flatbuffers by default

* fix hallucinated paths

* ensure imports present

* fix flatc target order

* fix elif syntax

* more rearrange

* fix git tag

* update template to allow proper configuration generation

* remove guard

* change quotes

* change quotes

* test default

* fix missing functions

* rearrange dependencies

* set the engine

* reintegrate some old cuda logic

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* Add back in function defs

* update api to use proper cuda versions

* update api to use proper cuda versions

* update api to use proper cuda versions

* Add back in function defs

* delete old paths

* aDd back include_ops.h generation

* delete old paths

* reinroduce old comand

* delete old paths

* delete old paths

* better tar unpacking

* fix for suffixes

* Add missing cpu sources

* fix duplicate profile

* change mac image

* change mac image

* get rid of old gpg key step

* get rid of javadoc on mac

* remove unneeded deps for mac

* update arm compute path

* remove old command checks

* remove gpg key step

* clang specific type erro

* clang specific type erro

* add another missing cpu path

* add another missing cpu path

* fix helpers sources ordering

* fix helpers sources ordering

* Add back createFromDescriptor

* fix helpers sources ordering

* fix helpers sources ordering

* fix helpers sources ordering

* fix helpers sources ordering

* Add back createFromDescriptor

* add debug for onednn

* add debug for onednn

* SFINAE for type aliases

* SFINAE for type aliases

* SFINAE for type aliases

* SFINAE for type aliases

* ensure we tell compiler we're ok with certain apple version minimums

* ensure we tell compiler we're ok with certain apple version minimums

* ensure we tell compiler we're ok with certain apple version minimums

* ensure we tell compiler we're ok with certain apple version minimums

* fix imports

* update linker path

* fix lock type usage

* change order of sources

* share mutex types

* share mutex types

* share mutex types

* update linker path

* share mutex types

* share mutex types

* share mutex types

* decrease type pairs for sort

* decrease type pairs for sort

* decrease type pairs for sort

* decrease type pairs for sort

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* update the special methods to use combinations

* standardize output paths

* fix template paths

* refactor compiler flags

* update onednn to use similar approach to armcompute

* refactor compiler flags

* fix paths

* fix paths

* fix paths

* fix paths

* fix paths

* fix paths

* fix paths

* fix paths

* change target expand types

* change target expand types

* add new pairwise types

* add new pairwise types

* add new pairwise types

* update linker paths

* add new pairwise types

* add new pairwise types

* add new pairwise types

* fix strings with transform

* fix strings with transform

* update linker paths

* update default values for libnd4j.outputPath

* update default values for libnd4j.outputPath

* update default values for libnd4j.outputPath

* update pom.xml namespaces

* Update .github/workflows/build-deploy-linux-cuda-12.6.yml

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Update .github/workflows/build-deploy-linux-x86_64.yml

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-05-08 04:52:38 +09:00
Adam Gibson 0b8805517e Refactoring lambdas (#10202)
Refactoring ndarray operators to make them less error prone
More conversion of ndarray references to pointers
Add scalar ops smoke tests
2025-03-21 17:33:32 +09:00
Adam Gibson 2f1d459933 Updates java cpp versions to 1.5.11 (cuda, openblas etc (#10188)
* javacpp version upgrades
add new cuda versions
address a few leftover api changes for nativeops
address new python api change

* remove commented code
2025-02-19 15:54:41 +09:00
Adam Gibson 47ffcebfff Add new offset usage for determining offsets removing indexOffset usage. (#10129)
* Add new offset usage for determining offsets removing indexOffset usage.

* more pom updates
2024-11-06 16:28:11 +09:00
Adam Gibson cc582096c2 Update javacpp versions (#10086)
* Update javacpp versions

* Fix cpython version
2024-08-05 18:16:13 +09:00
Adam Gibson 7051233bb4 Clean up pom.xml across modules: logging (#10075)
Remove old test in deeplearning4j-common-tests
Fix shape of created arrays in iris utils
2024-08-02 16:41:58 +09:00
Adam Gibson 603cf98dec Post merge follow up (openblas upgrades, misc test fixes) (#10002)
* Misc updates

* More test fixes
2023-06-10 05:31:55 +09:00
Adam Gibson b8e9d4f157 Add new dot product attention keras layer (#9948)
* Add dot product attention v2 op
Fix up matrix_band op to allow minLower/upper negative arguments
Add new contextNumInputs/Outputs for verifying the number of inputs/outputs in a
context pointer
Misc spacing/comments clean up

* Add dot product attention v2 op
Fix up matrix_band op to allow minLower/upper negative arguments
Add new contextNumInputs/Outputs for verifying the number of inputs/outputs in a
context pointer
Misc spacing/comments clean up
Clean up attention layer signature

* Add dot product attention v2 op
Fix up matrix_band op to allow minLower/upper negative arguments
Add new contextNumInputs/Outputs for verifying the number of inputs/outputs in a
context pointer
Misc spacing/comments clean up
Clean up attention layer signature

* Add dot product attention v2 op
Fix up matrix_band op to allow minLower/upper negative arguments
Add new contextNumInputs/Outputs for verifying the number of inputs/outputs in a
context pointer
Misc spacing/comments clean up
Clean up attention layer signature
Fix long data type casting

* Add dot product attention v2 op
Fix up matrix_band op to allow minLower/upper negative arguments
Add new contextNumInputs/Outputs for verifying the number of inputs/outputs in a
context pointer
Misc spacing/comments clean up
Clean up attention layer signature
Fix long data type casting

* Fix best guess loss variables. Prioritize finding external errors when creating samediff layers/vertices.
Add more exceptions to inf for situations where they are set on purpose like global pooling

* Handle rank 2 training

* Remove hard coded path

* Fix nits
2023-05-08 16:10:15 +09:00
Adam Gibson f8bb11ddf4 Fix undefined symbols, Migrate more ints -> sd::LongType (#9963)
* Remove extra op context changes due to performance regressions
Fix javacpp hack removals due to unsatisfied link on mac
Remove extra nd4j-profiler dependency in tests

* Fix up unsatisfied links by cleaning up definitions
Add new compiler flags support for nd4j-native
Add validation for missing method definitions
More clean up on ints/unsigned values to sd::Longs

* More cuda int -> sd::LongType changes

* Fix cpu builds

* Fix cpu/cuda linker flags

* Java side int -> long changes

* Fix debugging left overs, jackson versions

* Fix jackson versions in tests
2023-04-21 16:51:12 +09:00
Adam Gibson 8ccc121942 Update import mappings for release (#9949) 2023-04-10 05:29:47 +09:00
Adam Gibson 1d82ab71d7 Softmax bp optimization (#9944)
* Optimize softmax by removing an extra softmax pass and instead adding softmax the softmax input to the backward pass.

* Update ondnn platform implementation
2023-03-30 14:31:51 +09:00
Adam Gibson ee969e66f1 NLP fixes 2 ( better batching) (#9852)
* Update skipgram/cbow batching to direct to proper batch code path

* Reapply batch fix

* Reapply batch fix
Remove execution during iterateSample (should only happen during finish()

* Propagate workers properly
Optimize loops for hsoftmax

* Add worker/vector calculation separation. Allows configuration of number of vector calculation threads
Remove unused batchsequences fields

* Update workers for sequencers

* Remove unused inference code
Consolidate everything to efficient batching (revision upon benchmarking)
Clean up unused parameters on old code
Add new iteration option for skipgram_inference
# Conflicts:
#	deeplearning4j/deeplearning4j-nlp-parent/deeplearning4j-nlp/src/main/java/org/deeplearning4j/models/embeddings/learning/impl/elements/CBOW.java
#	deeplearning4j/deeplearning4j-nlp-parent/deeplearning4j-nlp/src/main/java/org/deeplearning4j/models/embeddings/learning/impl/elements/SkipGram.java

* Remove unused inference code
Consolidate everything to efficient batching (revision upon benchmarking)
Clean up unused parameters on old code
Add new iteration option for skipgram_inference
# Conflicts:
#	deeplearning4j/deeplearning4j-nlp-parent/deeplearning4j-nlp/src/main/java/org/deeplearning4j/models/embeddings/learning/impl/elements/CBOW.java
#	deeplearning4j/deeplearning4j-nlp-parent/deeplearning4j-nlp/src/main/java/org/deeplearning4j/models/embeddings/learning/impl/elements/SkipGram.java

Fix up batch cases to have equivalent accuracy and performance to older releases
Move iterations down to c++ from java
Clean up misc compilation errors
Upgrade kotlin version

* Remove unused inference code
Consolidate everything to efficient batching (revision upon benchmarking)
Clean up unused parameters on old code
Add new iteration option for skipgram_inference
# Conflicts:
#	deeplearning4j/deeplearning4j-nlp-parent/deeplearning4j-nlp/src/main/java/org/deeplearning4j/models/embeddings/learning/impl/elements/CBOW.java
#	deeplearning4j/deeplearning4j-nlp-parent/deeplearning4j-nlp/src/main/java/org/deeplearning4j/models/embeddings/learning/impl/elements/SkipGram.java

Fix up batch cases to have equivalent accuracy and performance to older releases
Move iterations down to c++ from java
Clean up misc compilation errors
Upgrade kotlin version
Fix training speed

* Remove unused inference code
Consolidate everything to efficient batching (revision upon benchmarking)
Clean up unused parameters on old code
Add new iteration option for skipgram_inference
# Conflicts:
#	deeplearning4j/deeplearning4j-nlp-parent/deeplearning4j-nlp/src/main/java/org/deeplearning4j/models/embeddings/learning/impl/elements/CBOW.java
#	deeplearning4j/deeplearning4j-nlp-parent/deeplearning4j-nlp/src/main/java/org/deeplearning4j/models/embeddings/learning/impl/elements/SkipGram.java

Fix up batch cases to have equivalent accuracy and performance to older releases
Move iterations down to c++ from java
Clean up misc compilation errors
Upgrade kotlin version
Fix training speed
Get rid of old print statements
Remove redundant data type checks with putScalar
Remove extra allocation of arrays when running batch case
2022-11-28 16:29:16 +09:00
Adam Gibson 0e19286047 Update openblas definitions for upgrade, migrate blas lapack codegen to mainline (#9843) 2022-11-04 17:21:47 +09:00
Adam Gibson b8a1207576 Increase performance of skipgram/cbow allocation on batch size 1 (#9827)
* Fix https://github.com/deeplearning4j/deeplearning4j/issues/9112
Optimize allocation of non samediff ops
Ensure we use detach knob correctly in array cache memory mgr
Update graalvm configs for samediff namespaces
Remove commented code in various places
Add cache for temp pointers
Clean up imports
Directly optimize skipgram to use byte codes rather than ints in softmax to avoid a type cast
In skipgram since we use workspaces in the vector calculation threads, these should be attached to that workspace for the given thread

* Fix https://github.com/deeplearning4j/deeplearning4j/issues/9112
Optimize allocation of non samediff ops
Ensure we use detach knob correctly in array cache memory mgr
Update graalvm configs for samediff namespaces
Remove commented code in various places
Add cache for temp pointers
Clean up imports
Directly optimize skipgram to use byte codes rather than ints in softmax to avoid a type cast
In skipgram since we use workspaces in the vector calculation threads, these should be attached to that workspace for the given thread
Optimize allocation for cbow, adding new _inference focused op that's faster on single batch sizes by passing ints instead of arrays

* Fix https://github.com/deeplearning4j/deeplearning4j/issues/9112
Optimize allocation of non samediff ops
Ensure we use detach knob correctly in array cache memory mgr
Update graalvm configs for samediff namespaces
Remove commented code in various places
Add cache for temp pointers
Clean up imports
Directly optimize skipgram to use byte codes rather than ints in softmax to avoid a type cast
In skipgram since we use workspaces in the vector calculation threads, these should be attached to that workspace for the given thread
Optimize allocation for cbow, adding new _inference focused op that's faster on single batch sizes by passing ints instead of arrays
Add updated op descriptor for tests
2022-10-28 16:18:21 +09:00
Adam Gibson 635f749f8b Fixes #9826 (#9829) 2022-10-28 16:18:11 +09:00
Adam Gibson 3f0793e568 Fix standardize to avoid divide by zero (#9822)
* Fix standardize to avoid divide by zero
Clean up commented old code present in standardize

* Fix codegen module declarations

* Remove debug message
2022-10-23 22:27:58 +09:00
Adam Gibson 8b0996203d Migrate codegen modules from contrib to mainline so they can be used in other projects (#9820)
Sync versions in codegen modules as part of migration
Upgrade all kotlin versions to be inline with kotlin 1.7.20
Fix compiler errors related to missing cases in various classes due to compiler upgrade
2022-10-23 06:47:43 +09:00