This is the list of papers you need to read for each course topic. You will need to submit a weekly paper review for a subset of the papers, listed here.
Intro:
[1] A. J. Smith, "The Task of the Referee," IEEE Computer, 1990.
[2] N. Binkert, B. Beckmann, G. Black, S. K. Reinhardt, A. Saidi, A. Basu, J. Hestness, D. R. Hower, T. Krishna, S. Sardashti, R. Sen, K. Sewell, M. Shoaib, N. Vaish, M. D. Hill and D. A. Wood, "The gem5 Simulator," ACM SIGARCH Computer Architecture News, 2011.
[3] T. Nowatzki, J. Menon, C.-H. Ko and K. Sankaralingam, "Architectural Simulators Considered Harmful," IEEE Micro, 2015.
Instruction Set Architecture:
[4] E. Blem, J. Menon and K. Sankaralingam, "Power Struggles: Revisiting the RISC vs. CISC Debate on Contemporary ARM and x86 Architectures," HPCA, 2013.
[5] A. Venkat, H. Basavaraj and D. M. Tullsen, "Composite-ISA Cores: Enabling Multi-ISA Heterogeneity Using a Single ISA," HPCA, 2019.
[6] J. L Gustafson and I. T. Yonemoto, "Beating Floating Point at its Own Game: Posit Arithmetic," SuperFri, 2017.
Pipelining:
[7] A. González, F. Latorre and G. Magklis, "Processor Microarchitecture: An Implementation Perspective," Synthesis Lectures on Computer Architecture, 2011, Chapter 1.
[8] V. Srinivasan, D. Brooks, M. Gschwind, P. Bose, V. Zyuban, P. N. Strenski and P. G. Emma, "Optimizing Pipelines for Power and Performance," MICRO, 2002.
[9] A. Tiwari, S. R. Sarangi and J. Torrellas, "ReCycle: Pipeline Adaptation to Tolerate Process Variation," ISCA, 2007.
Instruction Flow:
[7] A. González, F. Latorre and G. Magklis, "Processor Microarchitecture: An Implementation Perspective," Synthesis Lectures on Computer Architecture, 2011, Chapters 3-4.
[10] J. E. Smith and A. R. Pleszkun, "Implementing Precise Interrupts in Pipelined Processors," IEEE Transactions on Computers, 1988.
[11] R. Sheikh, J. Tuck and E. Rotenberg, "Control-Flow Decoupling," MICRO, 2012.
[12] D. A. Jimenez and C. Lin, "Dynamic Branch Prediction with Perceptrons," HPCA, 2001.
[13] J. Albericio, J. San Miguel, N. Enright Jerger and A. Moshovos, "Wormhole: Wisely Predicting Multidimensional Branches," MICRO, 2014.
Register Data Flow:
[7] A. González, F. Latorre and G. Magklis, "Processor Microarchitecture: An Implementation Perspective," Synthesis Lectures on Computer Architecture, 2011, Chapters 5, 6.1-6.3, 7, 8.
[14] G. S. Sohi and S. Vajapeyam. "Instruction Issue Logic for High-Performance, Interruptable Pipelined Processors," ISCA, 1987.
[15] M. H. Lipasti and J. P. Shen, "Exceeding the Dataflow Limit via Value Prediction," MICRO, 1996.
[16] H. Tabani, J.-M. Arnau, J. Tubella and A. González, "A Novel Register Renaming Technique for Out-of-Order Processors," HPCA, 2018.
[17] T. Koizumi, R. Shioya, S. Sugita, T. Amano, Y. Degawa, J. Kadomoto, H. Irie and S. Sakai, "Clockhands: Rename-free Instruction Set Architecture for Out-of-order Processors," MICRO, 2023.
Memory Data Flow:
[7] A. González, F. Latorre and G. Magklis, "Processor Microarchitecture: An Implementation Perspective," Synthesis Lectures on Computer Architecture, 2011, Chapters 2, 6.4-6.5.
[18] A. Moshovos, S. E. Breach, T. N. Vijaykumar and G. S. Sohi, "Dynamic Speculation and Synchronization of Data Dependences," ISCA, 1997.
[19] T. E. Carlson, W. Heirman, O. Allam, S. Kaxiras and L. Eeckhout, "The Load Slice Core Microarchitecture," ISCA, 2015.
Cache and Memory Architecture:
[20] B. Jacob, "The Memory System: You Can't Avoid It, You Can't Ignore It, You Can't Fake It," Synthesis Lectures on Computer Architecture, 2009, Chapters 1-3.
[21] D. Sanchez and C. Kozyrakis, "The ZCache: Decoupling Ways and Associativity," MICRO, 2010.
[22] B. Falsafi and T. F. Wenisch, "A Primer on Hardware Prefetching," Synthesis Lectures on Computer Architecture, 2014, Chapters 2-3.
[23] X. Yu, C. J. Hughes, N. Satish and S. Devadas, "IMP: Indirect Memory Prefetcher," MICRO, 2015.
[24] S. Sardashti, A. Arelakis, P. Stenstrom and D. A. Wood, "A Primer on Compression in the Memory Hierarchy," Synthesis Lectures on Computer Architecture, 2015, Chapters 2-4.
[25] A. Ghasemazar, P. Nair and M. Lis, "Thesaurus: Efficient Cache Compression via Dynamic Clustering," ASPLOS, 2020.
[26] E. Choukse, M. Erez and A. R. Alameldeen, "Compresso: Pragmatic Main Memory Compression," MICRO, 2018.
Virtual Memory:
[27] B. Pham, V. Vaidyanathan, A. Jaleel and A. Bhattacharjee, "CoLT: Coalesced Large-Reach TLBs," MICRO, 2012.
[28] A. Sembrant, E. Hagersten and D. Black-Shaffer, "TLC: A Tag-Less Cache for Reducing Dynamic First Level Cache Energy," MICRO, 2013.
[29] D. Skarlatos, A. Kokolis, T. Xu and J. Torrellas, "Elastic Cuckoo Page Tables: Rethinking Virtual Memory Translation for Parallelism," ASPLOS, 2020.
Security and Reliability:
[30] M. K. Qureshi, "CEASER: Mitigating Conflict-Based Cache Attacks via Encrypted-Address and Remapping," MICRO, 2018.
[31] S. Deng, W. Xiong and J. Szefer, "A Benchmark Suite for Evaluating Caches' Vulnerability to Timing Attacks," ASPLOS, 2020.
[32] D. Dangwal, W. Cui, J. McMahan and T. Sherwood, "Safer Program Behavior Sharing through Trace Wringing," ASPLOS, 2019.
[33] Y. Kim, R. Daly, J. Kim, C. Fallin, J. H. Lee, D. Lee, C. Wilkerson, K. Lai and O. Mutlu, "Flipping Bits in Memory Without Accessing Them: An Experimental Study of DRAM Disturbance Errors," ISCA, 2014.
[34] Z. Deng, A. Feldman, S. A. Kurtz and F. T. Chong, "Lemonade from Lemons: Harnessing Device Wearout to Create Limited-Use Security Architectures," ISCA, 2017.
Non-Traditional Computing:
[35] M. Hicks, "Clank: Architectural Support for Intermittent Computation," ISCA, 2017.
[36] A. Bhattacharyya, A. Somashekhar and J. San Miguel, "NvMR: Non-Volatile Memory Renaming for Intermittent Computing," ISCA, 2022.
[37] J. San Miguel, M. Badr and N. Enright Jerger, "Load Value Approximation," MICRO, 2014.
[38] I. Akturk and U. R. Karpuzcu, "AMNESIAC: Amnesic Automatic Computer," ASPLOS, 2017.
[39] A. Madhavan, T. Sherwood and D. Strukov, "Race Logic: A Hardware Acceleration for Dynamic Programming Algorithms," ISCA, 2014.