TailClipper: Reducing Tail Response Time of Distributed Services Through System-Wide Scheduling | Proceedings of the 2024 ACM Symposium on Cloud Computing

· ACM Conferences

10 min read Original article ↗

Abstract

Abstract

Reducing tail latency has become a crucial issue for optimizing the performance of online cloud services and distributed applications. In distributed applications, there are many causes of high end-to-end tail latency, including operating system delays, request re-ordering due to fan-out/fanin, and network congestion. Although recent research has focused on reducing tail latency for individual application components, such as by replicating requests and scheduling, in this paper, we argue for a holistic approach for reducing the end-to-end tail latency across application components. We propose TailClipper, a distributed scheduler that tags each arriving request with an arrival timestamp, and propagates it across the microservices' call chain. TailClipper then uses arrival timestamps to implement an oldest request first scheduler that combines global first-come first serve with a limited form of processor sharing to reduce end-to-end tail latency. In doing so, TailClipper can counter the performance degradation caused by request reordering in multi-tiered and microservices-based applications. We implement TailClipper as a userspace Linux scheduler and evaluate it using cloud workload traces and a real-world microservices application. Compared to state-of-the-art schedulers, our experiments reveal that TailClipper improves the 99th percentile response time by up to 81%, while also improving the mean response time and the system throughput by up to 54% and 29% respectively under high loads.

AI Summary

AI-Generated Summary (Experimental)

This summary was generated using automated tools and was not authored or reviewed by the article's author(s). It is provided to support discovery, help readers assess relevance, and assist readers from adjacent research areas in understanding the work. It is intended to complement the author-supplied abstract, which remains the primary summary of the paper. The full article remains the authoritative version of record. Click here to learn more.

Click here to comment on the accuracy, clarity, and usefulness of this summary. Doing so will help inform refinements and future regenerated versions.

To view this AI-generated plain language summary, you must have Premium access.

Formats available

You can view the full content in the following formats:

References

[1]

Adam Belay, George Prekas, Mia Primorac, Ana Klimovic, Samuel Grossman, Christos Kozyrakis, and Edouard Bugnion. 2017. The IX Operating System: Combining Low Latency, High Throughput, and Efficiency in a Protected Dataplane. ACM Transactions on Computer Systems (TOCS) 34, 4 (2017), 11. https://doi.org/10.1145/2997641

[2]

Mike Belshe. 2010. More Bandwidth Doesn't Matter (Much). https://bit.ly/3RUykxs.

[3]

Daniel S Berger, Benjamin Berg, Timothy Zhu, Siddhartha Sen, and Mor Harchol-Balter. 2018. RobinHood: Tail Latency Aware Caching-Dynamic Reallocation from Cache-Rich to Cache-Poor. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 195--212. https://doi.org/10.5555/3291168.3291183

[4]

Peter Bodík, Ishai Menache, Joseph (Seffi) Naor, and Jonathan Yaniv. 2014. Brief Announcement: Deadline-Aware Scheduling of Big-Data Processing Jobs. In SPAA. https://www.microsoft.com/en-us/research/publication/brief-announcement-deadline-aware-scheduling-of-big-data-processing-jobs/

[5]

Gunter Bolch, Stefan Greiner, Hermann de Meer, and Kishor S. Trivedi. 1998. Queueing networks and Markov chains: modeling and performance evaluation with computer science applications. Wiley-Interscience, USA.

[6]

Eric Boutin, Jaliya Ekanayake, Wei Lin, Bing Shi, Jingren Zhou, Zhengping Qian, Ming Wu, and Lidong Zhou. 2014. Apollo: scalable and coordinated scheduling for cloud-scale computing. In Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation (Broomfield, CO) (OSDI'14). USENIX Association, USA, 285--300.

[7]

Gary Bradski, Adrian Kaehler, et al. 2000. OpenCV. Dr. Dobb's journal of software tools 3, 2 (2000).

[8]

Gary Bradski, Adrian Kaehler, et al. 2024. OpenCV - Smoothing Images. https://docs.opencv.Org/4.x/d4/d13/tutorial_py_filtering.html.

[9]

Nathan Bronson, Abutalib Aghayev, Aleksey Charapko, and Timothy Zhu. 2021. Metastable Failures in Distributed Systems. In Proceedings of the Workshop on Hot Topics in Operating Systems. 221--227. https://doi.org/10.1145/3458336.3465286

[10]

Adrian M Caulfield, Eric S Chung, Andrew Putnam, Hari Angepat, Jeremy Fowers, Michael Haselman, Stephen Heil, Matt Humphrey, Puneet Kaur, Joo-Young Kim, et al. 2016. A Cloud-Scale Acceleration Architecture. In The 49th Annual IEEE/ACM International Symposium on Microarchitecture. IEEE Press, 7. https://doi.org/10.1109/MICRO.2016.7783710

[11]

Jeffrey Dean and Luiz André Barroso. 2013. The Tail at Scale. Commun. ACM 56, 2 (Feb. 2013). https://doi.org/10.1145/2408776.2408794

[12]

Giuseppe DeCandia, Deniz Hastorun, Madan Jampani, Gunavardhan Kakulapati, Avinash Lakshman, Alex Pilchin, Swaminathan Sivasubramanian, Peter Vosshall, and Werner Vogels. 2007. Dynamo: Amazon's Highly Available Key-value Store. SIGOPS Oper. Syst. Rev. 41, 6 (Oct. 2007). https://doi.org/10.1145/1323293.1294281

[13]

Aldric Degorre and Oded Maler. 2008. On scheduling policies for streams of structured jobs. In International Conference on Formal Modeling and Analysis of Timed Systems. Springer, 141--154.

[14]

David Desmeurs, Cristian Klein, Alessandro Vittorio Papadopoulos, and Johan Tordsson. 2015. Event-Driven Application Brownout: Reconciling High Utilization and Low Tail Response Times. In Cloud and Autonomic Computing (ICCAC). https://doi.org/10.1109/ICCAC.2015.25

[15]

Ahmed Eleliemy and Florina M Ciorba. 2021. A Resourceful Coordination Approach for Multilevel Scheduling. arXiv preprint arXiv:2103.05809 (2021).

[16]

D. R. Engler, M. F. Kaashoek, and J. O'Toole. 1995. Exokernel: an operating system architecture for application-level resource management. In Proceedings of the Fifteenth ACM Symposium on Operating Systems Principles (Copper Mountain, Colorado, USA) (SOSP '95). Association for Computing Machinery, New York, NY, USA, 251--266. https://doi.org/10.1145/224056.224076

[17]

Brad Fitzpatrick. 2004. Distributed Caching with Memcached. Linux journal 2004, 124 (2004), 5.

[18]

Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, and Adam Belay. 2020. Caladan: Mitigating Interference at Microsecond Timescales. In Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation. 281--297. https://doi.org/10.5555/3488766.3488782

[19]

Kristen Gardner, Samuel Zbarsky, Sherwin Doroudi, Mor Harchol-Balter, Esa Hyytiä, and Alan Scheller-Wolf. 2016. Queueing with Redundant Requests: Exact Analysis. Queueing Systems 83, 3 (2016). https://doi.org/10.1007/s11134-016-9485-y

[20]

Google. [n. d.]. GhOSt: Fast & Flexible User-Space Delegation of Linux Scheduling. https://github.com/google/ghost-userspace https://github.com/google/ghost-userspace.

[21]

Md E. Haque, Yong hun Eom, Yuxiong He, Sameh Elnikety, Ricardo Bianchini, and Kathryn S. McKinley. 2015. Few-to-Many: Incremental Parallelism for Reducing Tail Latency in Interactive Services. In Architectural Support for Programming Languages and Operating Systems (ASPLOS). ACM. https://doi.org/10.1145/2694344.2694384

[22]

Md E Haque, Yuxiong He, Sameh Elnikety, Thu D Nguyen, Ricardo Bianchini, and Kathryn S McKinley. 2017. Exploiting Heterogeneity for Tail Latency and Energy Efficiency. In Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture. ACM, 625--638. https://doi.org/10.1145/3123939.3123956

[23]

Jack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse, Barret Rhoden, Josh Don, Luigi Rizzo, Oleg Rombakh, Paul Turner, and Christos Kozyrakis. 2021. GhOSt: Fast & Flexible User-Space Delegation of Linux Scheduling. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles (Virtual Event, Germany) (SOSP '21). Association for Computing Machinery, New York, NY, USA, 588--604. https://doi.org/10.1145/3477132.3483542

[24]

Yangqing Jia and Evan Shelhamer. 2024. Caffe Model Zoo. http://caffe.berkeleyvision.org/model_zoo.

[25]

Kostis Kaffes, Timothy Chong, Jack Tigar Humphries, Adam Belay, David Mazieres, and Christos Kozyrakis. 2019. Shinjuku: Preemptive Scheduling for μsecond-scale Tail Latency. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). 345--360. https://doi.org/10.5555/3323234.3323264

[26]

Rishi Kapoor, George Porter, Malveeka Tewari, Geoffrey M Voelker, and Amin Vahdat. 2012. Chronos: Predictable Low Latency for Data Center Applications. In Proceedings of the Third ACM Symposium on Cloud Computing. ACM, 9. https://doi.org/10.1145/2391229.2391238

[27]

Harshad Kasture and Daniel Sanchez. 2014. Ubik: Efficient Cache Sharing with Strict QoS for Latency-Critical Workloads. In ACM SIGPLAN Notices, Vol. 49. ACM, 729--742. https://doi.org/10.1145/2644865.2541944

[28]

Jinhan Kim, Sameh Elnikety, Yuxiong He, Seung-won Hwang, and Shaolei Ren. 2013. QACO: Exploiting Partial Execution in Web Servers. In Cloud and Autonomic Computing Conference (CAC). ACM, Article 12. https://doi.org/10.1145/2494621.2494636

[29]

Ron Kohavi, Alex Deng, Brian Frasca, Toby Walker, Ya Xu, and Nils Pohlmann. 2013. Online Controlled Experiments at Large Scale. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining. 1168--1176. https://doi.org/10.1145/2487575.2488217

[30]

Jacob Leverich and Christos Kozyrakis. 2014. Reconciling High Server Utilization and Sub-millisecond Quality-of-service. In European Conference on Computer Systems (EuroSys). ACM, Article 4. https://doi.org/10.1145/2592798.2592821

[31]

Jialin Li, Naveen Kr. Sharma, Dan R. K. Ports, and Steven D. Gribble. 2014. Tales of the Tail: Hardware, OS, and Application-level Sources of Tail Latency. In Symposium on Cloud Computing (SoCC). ACM, Article 9. https://doi.org/10.1145/2670979.2670988

[32]

Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang, Yu Ding, Jian He, and Chengzhong Xu. 2021. Characterizing Microservice Dependency and Performance: Alibaba Trace Analysis. In Proceedings of the ACM Symposium on Cloud Computing. 412--426. https://doi.org/10.1145/3472883.3487003 https://github.com/alibaba/clusterdata.

[33]

Marissa Mayer. 2006. What Google Knows. Proceedings of the Third Annual Web 2 (2006).

[34]

Jayakrishnan Nair, Adam Wierman, and Bert Zwart. 2010. Tail-Robust Scheduling via Limited Processor Sharing. Performance Evaluation 67, 11 (2010), 978--995. https://doi.org/10.1016/j.peva.2010.08.012

[35]

Samuel S. Ogden, Xiangnan Kong, and Tian Guo. 2021. PieSlicer: Dynamically Improving Response Time for Cloud-based CNN Inference. In 12th ACM/SPEC International Conference on Performance Engineering. Association for Computing Machinery (ACM).

[36]

Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, and Hari Balakrishnan. 2019. Shenango: Achieving High CPU Efficiency for Latency-Sensitive Datacenter Workloads. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). 361--378. https://doi.org/10.5555/3323234.3323265

[37]

Chandandeep Pabla. 2009. Completely Fair Scheduler. Linux Magazine 184 (2009).

[38]

Simon Peter, Jialin Li, Irene Zhang, Dan RK Ports, Doug Woos, Arvind Krishnamurthy, Thomas Anderson, and Timothy Roscoe. 2016. Arrakis: The Operating System is the Control Plane. ACM Transactions on Computer Systems (TOCS) 33, 4 (2016), 11. https://doi.org/doi/10.1145/2812806

[39]

George Prekas, Marios Kogias, and Edouard Bugnion. 2017. ZygOS: Achieving Low Tail Latency for Microsecond-scale Networked Tasks. In Proceedings of the 26th Symposium on Operating Systems Principles. ACM, 325--341. https://doi.org/10.1145/3132747.3132780

[40]

Henry Qin, Qian Li, Jacqueline Speiser, Peter Kraft, and John Ousterhout. 2018. Arachne: Core-Aware Thread Management. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 145--160. https://doi.org/10.5555/3291168.3291180

[41]

Ziv Scully, Lucas van Kreveld, Onno Boxma, Jan-Pieter Dorsman, and Adam Wierman. 2020. Characterizing Policies with Optimal Response Time Tails under Heavy-Tailed Job Sizes. In Abstracts of the 2020 SIGMETRICS/Performance Joint International Conference on Measurement and Modeling of Computer Systems (Boston, MA, USA) (SIGMETRICS '20). Association for Computing Machinery, New York, NY, USA, 35--36. https://doi.org/10.1145/3393691.3394179

[42]

Lalith Suresh, Marco Canini, Stefan Schmid, and Anja Feldmann. 2015. C3: Cutting Tail Latency in Cloud Data Stores via Adaptive Replica Selection. In 12th USENIX Symposium on Networked Systems Design and Implementation (NSDI 15). 513--527. https://doi.org/10.5555/2789770.2789806

[43]

Don Towsley and François Baccelli. 1991. Comparisons of Service Disciplines in a Tandem Queueing Network with Real Time Constraints. Operations Research Letters 10, 1 (1991), 49--55. https://doi.org/10.1016/0167-6377(91)90086-5

[44]

Bhuvan Urgaonkar, Prashant Shenoy, Abhishek Chandra, and Pawan Goyal. 2005. Dynamic provisioning of multi-tier internet applications. In Second International Conference on Autonomic Computing (ICAC'05). IEEE, 217--228.

[45]

Balajee Vamanan, Hamza Bin Sohail, Jahangir Hasan, and TN Vijaykumar. 2015. Timetrader: Exploiting Latency Tail to Save Datacenter Energy for Online Search. In Proceedings of the 48th International Symposium on Microarchitecture. ACM, 585--597. https://doi.org/10.1145/2830772.2830779

[46]

Ashish Vulimiri, Philip Brighten Godfrey, Radhika Mittal, Justine Sherry, Sylvia Ratnasamy, and Scott Shenker. 2013. Low Latency via Redundancy. In Proceedings of the ninth ACM conference on Emerging networking experiments and technologies. ACM, 283--294. https://doi.org/10.1145/2535372.2535392

[47]

Qingyang Wang, Yasuhiko Kanemasa, Jack Li, Chien-An Lai, Chien-An Cho, Yuji Nomura, and Calton Pu. 2014. Lightning in the Cloud: A Study of Very Short Bottlenecks on n-Tier Web Application Performance. In Proceedings of USENIX Conference on Timely Results in Operating Systems. https://doi.org/10.13140/2.1.1479.0402

[48]

Christo Wilson, Hitesh Ballani, Thomas Karagiannis, and Ant Rowtron. 2011. Better never than late: meeting deadlines in datacenter networks. In Proceedings of the ACM SIGCOMM 2011 Conference (Toronto, Ontario, Canada) (SIGCOMM '11). Association for Computing Machinery, New York, NY, USA, 50--61. https://doi.org/10.1145/2018436.2018443

[49]

Zhe Wu, Curtis Yu, and Harsha V Madhyastha. 2015. CosTLO: Cost-Effective Redundancy for Lower Latency Variance on Cloud Storage Services. In 12th USENIX Symposium on Networked Systems Design and Implementation (NSDI 15). 543--557. https://doi.org/10.5555/2789770.2789808

[50]

Kenichi Yasukata, Michio Honda, Douglas Santry, and Lars Eggert. 2016. StackMap: Low-Latency Networking with the OS Stack and Dedicated NICs. In 2016 USENIX Annual Technical Conference (USENIX ATC 16). 43--56.

[51]

Matei Zaharia, Mosharaf Chowdhury, Michael J. Franklin, Scott Shenker, and Ion Stoica. 2010. Spark: cluster computing with working sets. In Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing (Boston, MA) (HotCloud'10). USENIX Association, USA, 10.

[52]

Zhizhou Zhang, Murali Krishna Ramanathan, Prithvi Raj, Abhishek Parwal, Timothy Sherwood, and Milind Chabbi. 2022. CRISP: Critical Path Analysis of Large-Scale Microservice Architectures. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). USENIX Association, 655--672. https://www.usenix.org/conference/atc22/presentation/zhang-zhizhou