Perhaps I'm sick in the mind but I really like debugging, especially hard ones.
Recently our stack fell over when we did a production release. Basically web worker threads were spawning more worker threads and this crashed our AOP proxy (Castle.Core) which uses some global state and was spinlocking for some reason. We ended up with everything whacked at 100% across the cluster. 2 hours into the problem and it was something completely different - a for loop that didn't end on another thread and took up one core per node causing CPU contention. The spinlocks just amplified the problem.
One change to the loop continuation clause and wham -- fixed.
The journey is exciting and you learn so much on the way. It's great.
Recently our stack fell over when we did a production release. Basically web worker threads were spawning more worker threads and this crashed our AOP proxy (Castle.Core) which uses some global state and was spinlocking for some reason. We ended up with everything whacked at 100% across the cluster. 2 hours into the problem and it was something completely different - a for loop that didn't end on another thread and took up one core per node causing CPU contention. The spinlocks just amplified the problem.
One change to the loop continuation clause and wham -- fixed.
The journey is exciting and you learn so much on the way. It's great.