autumn-2-net
76afe57e47
New generative model algorithm: Rectified Flow ( #184 )
...
Implements Rectified Flow in DiffSinger.
ref:
- https://github.com/gnobitab/RectifiedFlow
- https://github.com/yxlllc/ReFlow-VAE-SVC
---------
Co-authored-by: yqzhishen <yangqian_1015@icloud.com >
2024-04-17 22:57:54 +08:00
yqzhishen
71b8fbe5b6
Fix path type
2024-04-17 21:49:15 +08:00
yqzhishen
5ca6261deb
Refactor click options
2024-04-03 21:52:08 +08:00
yqzhishen
c16095bc55
Drop support for some old features and behaviors ( #172 )
...
* Drop support for discrete F0 embedding (reserved in ONNX exporter)
* Drop support for `interp_uv` configuration key
* Drop support for `train_set_name` and `valid_set_name` configuration keys
* Drop support for linear domain of random time stretching augmentation
* Drop support for `num_pad_tokens` configuration key
* Drop support for code backup before training
* Drop support for `ffn_padding` configuration key
* Drop support for random seeding
* Add placeholder to load old checkpoint
* Remove duplicate txt_embed layer (resuming may raise errors)
* Remove migration script and error message for transcriptions.txt
* Remove seed from batch shuffling
* Use direct access on some hparam keys
* Fix duplicate keys in YAML
* Rename `pndm_speedup` to `diff_speedup`
2024-02-25 20:57:02 +08:00
yqzhishen
f68df7bf94
Support specifying vocoder model path for exporting
2023-11-29 02:04:26 +08:00
yqzhishen
859ad2e3ec
Expose expr by default and support --freeze_glide
2023-11-25 11:24:35 +08:00
yqzhishen
fbff2e8376
Enable augmentation and expose parameters by default
2023-11-20 21:22:20 +08:00
Dachun Sun
36ac5f583b
Multi-node batched validation and improvement on strategy selection ( #148 )
...
* More metadata when binarize
* Picklable AttrDict
* Support other notion of 'sizes'
* Better strategy getter, Ds TB logger, and eval batch sampler.
* Add title to plots and specify figsize
* Unify batch sampler and bug fixes
* Batch and multi-device validation
* Fix imports
* Fix deadlock under multigpu and resume from ckpt
* Remove unnecessary if in base_dataset
* Move build loss to finish init and ft to build model
* Prevent repeated valid item
* Remove optimizer_idx to support lightning 2.1
* Warning message fix
* Fix val error when aux is off
* Rename fields in metadata and add doc
* Adjust respective duration logging for each speaker
* val persisent worker, module list ordering, update doc
---------
Co-authored-by: yqzhishen <yangqian_1015@icloud.com >
2023-10-29 00:37:14 +08:00
autumn-2-net
0ec98bdbf9
Shallow diffusion and aux decoder ( #128 )
...
* Add shallow diffusion API
* Support aux decoder training
* Support shallow diffusion inference
* add shallow farmwork
* add shallow farmwork
* Support lambda for aux mel loss
* Move config key
* add shallow farmework
* add shallow farmework
* add denorm
* add shallow model training switch
* Limit gradient from aux decoder
* Improve loss calculation control flow
* add independent encoder in shallow
* Adjust lambda
* Implement shallow diffusion
There are some issues to resolve in DPM-Solver++ and UniPC
* Fix missing depth assignment
* fix bugs of shallow diffusion inference
* Fix errors and remove debug code
* Support K_step < timesteps (shallow-only diffusion)
* Fix argument passing
* Add missing checks
* add glow decoder
* add glow decoder
* add convnext glow decoder
* fix fs2
* Support using gt mel as source during validation
* Clean files and configs
* Clean and refactor aux decoder
* Fix KeyError
* Support exporting shallow diffusion to ONNX
* Add missing logic to ONNX
* Rename `diff_depth` to `K_step_infer`
---------
Co-authored-by: yqzhishen <yangqian_1015@icloud.com >
Co-authored-by: autumn <2>
Co-authored-by: llc1995@sina.com <llc1995@sina.com >
2023-09-23 00:50:02 +08:00
yqzhishen
38bc407156
Implement pitch expressiveness controlling mechanism ( #97 )
...
* Add expressiveness in model `forward()`
* Support inference with static or dynamic expressiveness
* Fix assignment of `retake_`
* Format code
* Add `expressiveness` in ONNX model
* Swap input order
* Fix typo
* Adapt latest updates from main branch
2023-08-11 10:31:46 +08:00
yqzhishen
88d8ae5110
Fix wrong model loading logic when using --mel
2023-08-01 14:09:35 +08:00
yqzhishen
95e3e7d27a
Perform graceful exit on KeyboardInterrupt ( #119 )
2023-07-20 00:19:54 +08:00
autumn-2-net
7847af2d11
support finetuning from pretrained checkpoints ( #108 )
...
* support pre_train model add doc
* Update docs for finetuning
---------
Co-authored-by: autumn <2>
Co-authored-by: yqzhishen <yangqian_1015@icloud.com >
2023-07-17 21:24:22 +08:00
yqzhishen
a2e388dd71
Support variance models in drop_spk.py
2023-07-16 22:01:29 +08:00
yqzhishen
b658499ec5
Support spk mix in variance exporter
2023-07-15 21:10:40 +08:00
yqzhishen
8a818b269c
Update descriptions and logging
2023-07-09 20:36:29 +08:00
yqzhishen
94c0b9f240
Add PYTHONPATH envs in binarize.py and train.py
2023-06-13 20:49:30 +08:00
yqzhishen
51acdde675
Restore compatibility for Python 3.8
2023-06-13 20:45:18 +08:00
yqzhishen
af4d8ec8e6
Do not support old DS files anymore
2023-06-13 19:24:21 +08:00
yqzhishen
1ab32defca
Remove default value of --gender to allow None input
2023-06-13 00:43:39 +08:00
yqzhishen
4947f9b0dd
Fix typo
2023-06-12 23:48:15 +08:00
yqzhishen
c1208b81c3
Fix missing spk_mix in variance model inference
2023-06-02 23:29:41 +08:00
yqzhishen
b5eacb135d
Support speaker mix in variance model
2023-06-02 18:22:43 +08:00
yqzhishen
708ec58966
Support variance model inference from CLI
2023-06-02 00:15:26 +08:00
yqzhishen
5d0c348e62
Fix --speedup not working
2023-05-29 23:02:38 +08:00
yqzhishen
2de7a21f79
Refactor inference structure
2023-05-28 01:32:43 +08:00
yqzhishen
c7303ef30c
Finish variance model exporting
2023-05-21 13:52:18 +08:00
yqzhishen
c29e073d02
Fix label format to CSV and add migrating scripts
2023-05-18 18:28:22 +08:00
yqzhishen
3b0748029d
Support more argument formats
2023-04-11 23:09:03 +08:00
yqzhishen
8701b8955d
Add script to drop speaker embedding from checkpoints
2023-04-11 22:44:09 +08:00
yqzhishen
c88733b9ea
Migrate some path operations to pathlib
2023-04-11 13:18:05 +08:00
yqzhishen
d4943bba11
Merge pull request #75 from yqzhishen/refactor-onnx
...
Re-implement ONNX exporting scripts to fit new PyTorch versions
2023-04-11 01:50:33 +08:00
yqzhishen
6a6de253ef
Optimize spk export and freeze logic
...
- if there is only one speaker, freeze him/her by default
- if there are multiple speakers but no --freeze_spk and no --export_spk is set, export them all
2023-04-09 23:44:26 +08:00
yqzhishen
81826ae1bb
Adjust stdout
2023-04-09 13:16:50 +08:00
yqzhishen
c9a8b105e5
Update checkpoints loading for NSF-HiFiGAN
2023-04-08 19:20:57 +08:00
yqzhishen
493a80dd96
Finish NSFHiFiGANExporter
2023-04-08 18:15:01 +08:00
yqzhishen
b70626e983
Finish export.py for acoustic exporter
2023-04-08 01:13:23 +08:00
yqzhishen
664a98e038
Fix gender NoneType bug
2023-04-07 00:55:32 +08:00
yqzhishen
c3b8ac6aff
Merge pull request #72 from hrukalive/refactor-pl
...
Refactor to support PyTorch 2.0 and Lightning 2.0
2023-04-04 22:58:24 +08:00
yqzhishen
2db8751bca
Re-organize infer_utils
2023-03-30 16:58:55 +08:00
yqzhishen
e632dda343
Merge branch 'refactor-v2' into refactor-pl
2023-03-27 14:55:43 +08:00
yqzhishen
f1bef04d3e
Fix torch.load error on pure-CPU machines
2023-03-27 00:43:15 +08:00
hrukalive
c1ab92af68
Revert back some small changes for diff
2023-03-26 00:11:44 -05:00
hrukalive
1a2f2c9a0a
Fix for reviews
2023-03-25 23:39:04 -05:00
hrukalive
0543914f9d
Auto strategy choose, gloo backend by default
2023-03-25 22:36:44 -05:00
hrukalive
2bbc42b3b0
Add env for CUDNN API change, clean more codes
2023-03-25 11:18:37 -05:00
hrukalive
93e4627a90
Use pl rankzero utils to discriminate main proc
2023-03-25 11:18:36 -05:00
hrukalive
ae2946c8fa
Initial attempt to refactor lightning code
2023-03-25 11:18:36 -05:00
yqzhishen
984ae5fe3c
Use positional arguments in migrate.py
2023-03-21 00:44:05 +08:00
yqzhishen
e28819e6e3
Use mp.Pool instead of mp.Process
2023-03-19 14:16:50 +08:00