09cc78599a
We restrict each algorithm time loop to a pre-determined amount of time. Exceeding this limit will cause the algorithm to immediately terminate. This quickly becomes an issue when considering users running trainable models that have a long initialization period that exceeds the time loop maximum. This change provides a mechanism through which a long-running scheduled event is permitted to keep running and is permitted to avoid the time loop permitted by requesting additional time. Requests for additional time are limited according to a leaky bucket implementation whose parameters are set via the job's controls structure. The fundamental time unit for the algorithm is a single minute. Here's how it works. If a scheduled event takes longer than one full wall clock second then a request is made to the leaky bucket for one more minute. If the scheduled event continues to take more time, it will continue to request additional minutes. Each requested minute will prevent the algorithm's time loop check from terminating the algorithm. When the bucket is empty and no more minutes are available to be requested, a TimeoutException is thrown causing a cascade that ends in the algorithm's termination and status being flipped to RuntimeError. Additionally, this applies equally to ALL scheduled events. While some helpers were added with the naming of Train and TrainNow to the ScheduleManager, these methods don't do anything special and the infrastructure doesn't otherwise flag them as different, so this feature becomes part of the core Scheduled Event feature set. Further, the live scheduled events were not touched and are still pending further discussion regarding the value added by enforcing a time restriction when simulation time and wall clock time are equivalent. Fixes #3319