• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Muse deploys fix after scheduled tasks fail to recover from outages

    Muse says several users received reports that their scheduled tasks had been failing for a while after the outages.

    AW
    TS
    2 Sources, ,

    TLDR

    Muse says it deployed a fix to improve recovery of scheduled tasks and cron jobs. During recent system changes, outages left its inference service unresponsive and virtual machines unavailable for longer during deployments. Some tasks did not resume afterward. The team says it has strengthened monitoring and alerts and is auditing its recovery systems.

    Combined views

    19.8K

    2 Sources, first seen 3h ago

    Combined views

    19.8K

    2 Sources, first seen 3h ago

    248 likes
    3h ago
    first seen 3h ago
    248 likes
    22 comments
    36 saves
    18 reposts
    22 comments
    36 saves
    18 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @bigT_sheeshScheduled Tasks on @Muse - An update on reliability We deployed a fix to Muse to improve scheduled task and cron recovery and reliability especially, and have hardened our monitoring and alerting of the cron systems in the VMs. Over the last few days the Muse team has been deploying a few changes to: Improve the KV cache hit rate after increasing the context window size of the underlying model Improve our VM scheduling algorithm to reduce downtime during deployments Deploy a new authentication token format across the fleet for more fine-grained access. During these changes we had a few SEVs (outages) where the inference tier was unresponsive, and the VM scheduling resulted in VMs not being available for longer periods of time during deployment and being unresponsive to clients. We observed after we recovered from these outages that several crons and scheduled tasks failed to resume / heal properly when the outage was done, resulting in several users receiving reports from their Muse that the crons have been failing for a while. We apologise for this incident, and know how much this disrupts your flow especially since many users rely on muse to reliably deliver outcomes for them. We take these outages seriously, and are running a full audit of our recovery systems and tools, and are adding more dashboards and alerts so you can continue to receive the high quality service we want all users to have.3h
    @alexandr_wangRT @bigT_sheesh: Scheduled Tasks on @Muse - An update on reliability We deployed a fix to Muse to improve scheduled task and cron recovery…2h

    2 Sources

    @bigT_sheeshScheduled Tasks on @Muse - An update on reliability We deployed a fix to Muse to improve scheduled task and cron recovery and reliability especially, and have hardened our monitoring and alerting of the cron systems in the VMs. Over the last few days the Muse team has been deploying a few changes to: Improve the KV cache hit rate after increasing the context window size of the underlying model Improve our VM scheduling algorithm to reduce downtime during deployments Deploy a new authentication token format across the fleet for more fine-grained access. During these changes we had a few SEVs (outages) where the inference tier was unresponsive, and the VM scheduling resulted in VMs not being available for longer periods of time during deployment and being unresponsive to clients. We observed after we recovered from these outages that several crons and scheduled tasks failed to resume / heal properly when the outage was done, resulting in several users receiving reports from their Muse that the crons have been failing for a while. We apologise for this incident, and know how much this disrupts your flow especially since many users rely on muse to reliably deliver outcomes for them. We take these outages seriously, and are running a full audit of our recovery systems and tools, and are adding more dashboards and alerts so you can continue to receive the high quality service we want all users to have.3h
    @alexandr_wangRT @bigT_sheesh: Scheduled Tasks on @Muse - An update on reliability We deployed a fix to Muse to improve scheduled task and cron recovery…2h