07-129 · Freshman Immigration Course · ማስታወሻዎች

07-129 Research Notes

Every research note from Carnegie Mellon’s first-year computer science seminar, collected in one place — from Randy Pausch on time management to AI and autonomous robots. Search across all of them or jump to a set.

7 sets34 questions & lessons~29 min total readCarnegie Mellon · SCS · Fall 2026

Set 01 · Craft

Time Management

Five lessons from Randy Pausch

5 lessons2 min readRandy Pausch · lecture

Randy Pausch was a computer science professor at Carnegie Mellon University and a pioneer in virtual reality research. He delivered highly influential lectures on optimizing professional and personal workflows prior to his passing.

01Treating Time as the Ultimate Currency#

Pausch makes a compelling case that time matters infinitely more than financial resources. Money can be regained, but spent time is permanently lost.

"Time is the only commodity that matters."

02Strategic Planning#

He emphasized the necessity of documenting goals and scheduling them actively to prevent external factors from dictating productivity.

"Failing to plan is planning to fail."

03Navigating the Urgent Versus the Important#

Using a quadrant system, he highlights the operational flaw of prioritizing loud, urgent distractions over quiet, foundational goals.

"You have to do the important stuff before it becomes urgent."

04The Heavy Cost of Interruptions#

Workflow disruptions carry significant cognitive penalties. Regaining focus after minor interruptions depletes critical mental energy.

"You do not have to check your email every two minutes."

05The Core Purpose of Efficiency#

Working faster is fundamentally counterproductive if the extra time is simply filled with more work. True efficiency aims to buy back personal time.

"Nobody on their deathbed ever said, 'I wish I had spent more time at the office.'"
Back to all sets

Set 02 · Languages

Programming Languages

Paradigms & history

4 questions4 min read

01Moving from Punch Cards to Programming Languages#

In my introductory coursework, making a simple syntax error in Python is already frustrating, so imagining physical punch cards makes me appreciate how much smoother software development has become. Studying early batch processing helped me understand why this transition was necessary. The core reasons for the shift include:

  • High Physical Error Rates and Slow Feedback: Punch cards encoded data using physical holes. A single typing mistake meant manually re-punching the card on a keypunch machine or facing failed batch execution jobs that took hours to run.
  • Need for Human Readability: Early programmers had to write raw binary instructions or hardware mnemonics, which made code dense and hard to debug. High-level languages like FORTRAN allowed developers to write mathematical formulas intuitively.
  • Portability Across Hardware: Punch card decks and early assembly code were bound to specific machine architectures. High-level programming languages enabled code portability by allowing compilers to translate unified source code for different target machines.
  • Complexity of Expressing Algorithms: Expressing control structures like loops and conditionals with physical cards led to unstructured code organization. Higher-level languages provided structured control flow mechanisms that reduced logical errors.

02The Need for Hundreds of Programming Languages#

While Python handles many general-purpose tasks, different problems demand different trade-offs in software engineering. Observing distinct runtime requirements clarifies why a single language cannot serve all technological domains:

  • Domain Optimization: Different industries require specialized toolsets. C is tailored for operating system kernels, JavaScript dominates browser environments, Python excels in data science, and C++ handles high-performance graphics engines.
  • Diverse Programming Paradigms: Languages offer different mental frameworks for structuring code, such as imperative, object-oriented, functional, or event-driven paradigms, depending on how logic needs to be modeled.
  • Performance vs. Productivity Trade-offs: Interpreted languages with automatic memory management offer rapid development velocity, whereas compiled languages with manual memory control provide raw execution speed.
  • Target Platforms and Runtimes: A browser runtime needs lightweight asynchronous event handling, while embedded microcontrollers require predictable, low-level hardware access.

03Drawbacks of Modern Languages and Proposed Enhancements#

In my technical projects, I primarily utilize Python and JavaScript. During development, I frequently encounter specific language behaviors that create structural limitations or bugs.

Key Drawbacks

JavaScript:

  • Implicit Type Coercion: JavaScript automatically converts types during operations, leading to unintuitive bugs (e.g., evaluating "5" + 3 results in "53", whereas "5" - 3 results in 2).
  • Single-Threaded Main Loop: JavaScript relies on an event loop running on a single thread. Heavy CPU computations can block this thread and freeze the application.

Python:

  • Global Interpreter Lock (GIL): Python utilizes a GIL that prevents true parallel execution of multiple CPU-bound threads.
  • Execution Speed Limitations: Being dynamically typed and interpreted introduces significant runtime performance overhead compared to compiled languages.

Proposed Structural Changes

  • Native Strict Typing Mode for JavaScript: Implementing an opt-in runtime keyword (like "use strict types") to completely disable implicit type coercion.
  • Built-In Parallel Processing Engine for Python: Integrating native multi-threading architecture that bypasses standard GIL restrictions for purely CPU-heavy workloads.

04Designing a New Programming Language#

Understanding compiler mechanics reveals how highly structured the translation from plain text to machine logic must be. The standard pipeline includes:

  1. Define Syntax and Language Goals: Establishing keyword syntax, operator rules, variable scoping, and dynamic or static typing rules based on the problem domain.
  2. Build a Lexer (Tokenizer): Creating a lexical analyzer that reads source code character by character and groups them into meaningful tokens (keywords, literals, operators).
  3. Build a Parser (Abstract Syntax Tree): Converting the flat token stream into a hierarchical Abstract Syntax Tree (AST), ensuring the code follows grammar rules and operator precedence.
  4. Construct an Execution Engine: Implementing an execution backend, either as a tree-walking interpreter that processes the AST directly, or a transpiler/compiler that generates target machine code.

References

  1. Wikipedia. Computer programming in the punched card era.
  2. Pyatagouda, S. The Evolution of Programming Languages: From Punch Cards to Modern Code. Medium.
  3. Quantum Zeitgeist. Programming Languages: From Punch Cards to Python.
  4. ResearchGate. An Insight into Programming Paradigms and Their Programming Languages.
  5. IBM. What is a Compiler?
  6. Bailey, J. Fundamentals of JavaScript.
Back to all sets

Set 03 · Theory

Theory of Computation

Decidability & P versus NP

5 questions3 min read

01What is a decision problem?#

A decision problem is a computational question that yields a binary output: either "yes" or "no" for a given set of input values. Formally, it can be defined as determining whether a given input string belongs to a specific formal language. Standard examples include:

  • Primality Testing: Determining whether a given integer n is a prime number.
  • Graph Connectivity: Determining whether two specific vertices in a graph are connected by a path.

02What does it mean for a decision problem to be decidable?#

A decision problem is decidable (or computable) if an algorithm exists that can process any valid input and yield the correct "yes" or "no" answer in a finite number of steps. If no algorithm can guarantee termination with a correct answer for every input, the problem is classified as undecidable. The Halting Problem, formulated by Alan Turing, is a classic example of an undecidable problem.

03What is the class P? What is the class NP?#

Class P (Polynomial Time): The class of decision problems that can be solved by a deterministic Turing machine within a time bound proportional to a polynomial function of the input size, expressed as O(nk) for some constant k. Problems in P are considered computationally tractable.

Class NP (Nondeterministic Polynomial Time): The class of decision problems for which a candidate solution can be verified by a deterministic Turing machine in polynomial time. Alternatively, these are problems that can be solved in polynomial time using a nondeterministic Turing machine.

04What is the intuitive meaning of the "P versus NP" question?#

The intuitive meaning of the "P versus NP" question centers on whether every problem whose solution can be easily verified by a computer can also be easily solved by a computer from scratch.

  • If P = NP, finding a solution to a complex problem is inherently as easy as checking a proposed solution.
  • If P ≠ NP, verifying the correctness of a solution is fundamentally easier than discovering that solution independently.

05If you resolve the P versus NP question, how much richer will you be?#

Resolving the P versus NP question yields a cash award of $1,000,000 USD. This monetary prize was established by the Clay Mathematics Institute in May 2000 as one of the seven Millennium Prize Problems.

References

  • Clay Mathematics Institute: The P vs NP Problem
  • Wikipedia: Decision Problem
  • Wikipedia: P versus NP problem
  • NIST Dictionary of Algorithms and Data Structures: Decidable Problem
Back to all sets

Set 04 · Human-centered

Human–Computer Interaction

Useful, usable & tangible

5 questions2 min read

01What is human-computer interaction (HCI)?#

Human-computer interaction is the study of how people use technology and how to design systems that are both effective and accessible. Writing the functional backend for a system is distinct from designing an intuitive frontend architecture. The primary disciplines that contribute to HCI include computer science, cognitive psychology, behavioral science, and graphic design.

02Useful versus Usable Systems#

A system is useful when it provides the exact features required to achieve a specific goal. A system is usable when users can operate those features easily and efficiently. Utility and usability combine to determine overall success. An interface might follow all usability guidelines perfectly, but if it solves a problem the user does not have, it holds zero utility.

03Everyday Interface Failures#

A common interface failure is the lack of system status visibility—for instance, digital payment portals that process transactions but reload without displaying a confirmation message. This often lacks semantic HTML structure, degrading accessibility for screen readers. A structural improvement involves implementing clear, high-contrast confirmation banners that dynamically update post-transaction.

04Prototypes and Testing#

A prototype is a preliminary model of a product used to validate design logic before deploying backend code. Testing a prototype with end-users helps identify navigational friction early in the pipeline. Rather than gathering subjective aesthetic feedback ("Do you like this color?"), rigorous prototype testing measures functional success: Can the user locate critical inputs, recover from errors, and navigate system state changes efficiently?

05Tangible Interfaces#

Interfaces extend beyond screens. The KIBO Robot utilizes a tangible interface built from physical blocks, removing traditional digital screens from the programming loop. This provides high accessibility for demographics unable to utilize standard keyboards (e.g., young children). Analyzing the efficacy of such interfaces requires comparative heuristic testing against traditional digital inputs.

References

  • CMU Human-Computer Interaction Institute. About the HCII.
  • Nielsen, J. 10 Usability Heuristics for User Interface Design. Nielsen Norman Group.
  • W3C Web Accessibility Initiative (WAI). Introduction to Web Accessibility.
  • DevTech Research Group, Boston College. KIBO Robot.
Back to all sets

Set 05 · Systems

Distributed Systems

When one computer is not enough

5 questions4 min readProf. Hammoud

01What happens when the problem you want to solve becomes too big for any one computer?#

When a problem becomes too massive for a single computer to handle, its physical hardware running out of memory or processing capacity causes the system to crash or slow down. To solve this, computer scientists use horizontal scaling, which means distributing the workload across a cluster of multiple computers connected over a network. A classic example is Google's MapReduce framework, which splits giant datasets into smaller chunks and hands them to individual worker machines, which process the data and send back partial answers.

Uncertainty Factor: It is difficult to predict exactly how well a given task will scale because some software algorithms cannot be cleanly broken down, and unexpected network bottlenecks can unpredictably stall the entire cluster.

02Suppose 1,000 computers work together. Do you now have one computer that is 1,000 times more powerful? Why or why not?#

A cluster of 1,000 computers does not equal one machine that is 1,000 times faster. In 1967, computer architect Gene Amdahl formulated Amdahl's Law, showing that every computer task contains portions that must run sequentially and cannot be split among parallel processors. Furthermore, coordinating 1,000 separate machines introduces massive communication overhead. The computers must pass data across network cables, coordinate their timing, deal with hardware bottlenecks, and wait for slower machines to finish. Think of it like a group of 1,000 people writing a single essay together, where people spend significant energy chatting and coordinating instead of writing.

Uncertainty Factor: Real-world performance speedups fluctuate depending on network congestion, hardware age differences, software efficiency, and unexpected background tasks, making exact speed calculations uncertain before running the code.

03Can 1,000 computers agree on something if some of them fail or even lie?#

Yes, 1,000 computers can reach an agreement even if some fail or send fake data, provided the system uses a Byzantine Fault Tolerant protocol. Computer scientist Leslie Lamport and his research team proved in 1982 that as long as honest computers make up more than two-thirds of the network, consensus remains achievable. Mathematically, if m is the number of failing or lying computers, the network needs at least 3m + 1 total computers to stay safe. Consensus systems like Paxos or Proof-of-Work use voting rounds and validation rules to filter out false statements.

Uncertainty Factor: If a sudden network partition isolates critical nodes or if bad actors secretly control over one-third of the total machines at once, consensus breaks down completely and the network can accept false information.

04When you use ChatGPT, Google, Instagram, or an online game, where is the computation actually happening?#

The computation is split between your personal device and giant remote data centers across the world. Your phone or laptop acts as a client that handles displaying graphics, recording user input, playing local audio, and managing user interface animations. Meanwhile, the heavy processing tasks, such as executing multi-billion-parameter AI models, processing web searches, running complex game physics, and hosting user databases, occur inside server farms containing thousands of interconnected computers.

Uncertainty Factor: Because cloud platforms dynamically balance server workloads across continents to cut costs and latency, identifying the exact building or rack of servers processing your request at any given second is nearly impossible.

05If you could make millions of computers behave like one dependable machine, what could humanity build that we cannot build today?#

If millions of computers worked seamlessly as one machine, humanity could build tools that require unfathomable processing power. For instance, we could create digital twins of planet Earth to simulate global climate patterns with extreme accuracy weeks in advance. Scientists could run real-time atomic biophysics models to design custom medical drugs in seconds. Engineers could build universal traffic and supply chain optimization grids to eliminate traffic jams. Astronomy teams could construct planetary-scale radio telescope networks to process massive cosmic signals instantly.

Uncertainty Factor: It remains uncertain whether physical limitations, such as the speed of light delaying messages across continents, will allow millions of distant machines to synchronize smoothly without introducing intolerable communication delays.

References

  • Amdahl, G. M. (1967). Validity of the single processor approach to achieving large scale computing capabilities. Proceedings of the AFIPS Spring Joint Computer Conference, 483–485. https://doi.org/10.1145/1465482.1465560
  • Armbrust, M., Fox, A., Griffith, R., Joseph, A., Katz, R., Konwinski, A., Lee, G., Patterson, D., Rabkin, A., Stoica, I., & Zaharia, M. (2010). A view of cloud computing. Communications of the ACM, 53(4), 50–58. https://doi.org/10.1145/1721654.1721672
  • Dean, J., & Ghemawat, S. (2004). MapReduce: Simplified data processing on large clusters. Sixth Symposium on Operating System Design and Implementation (OSDI), 137–150.
  • Lamport, L., Shostak, R., & Pease, M. (1982). The Byzantine Generals Problem. ACM Transactions on Programming Languages and Systems, 4(3), 382–401. https://doi.org/10.1145/357172.357176
Back to all sets

Set 06 · AI

Multimodal Learning & Embodied AI

Many senses, one body

5 questions6 min readProf. Bilal Taha · Sep 29, 2026

01What is a modality in AI? Give three examples of different modalities.#

A modality is a distinct type or channel of information through which data is represented and perceived. Each modality has its own structure and statistical properties: text is a sequence of discrete tokens, an image is a grid of pixel values, and audio is a continuous waveform over time. Because of these differences, each modality traditionally required its own specialized model. Three examples of different modalities are:

  • Vision: Images and video captured by cameras, represented as arrays of pixel intensities.
  • Language: Written or transcribed text, represented as sequences of words or tokens.
  • Audio: Speech, music, and environmental sounds, represented as sound waves or spectrograms.

Other modalities include depth maps, LiDAR point clouds, touch (haptic) signals, and physiological signals such as heart rate.

Uncertainty Factor: The boundary between modalities is not always sharp. For example, a photo of a printed page contains text inside an image, and researchers disagree on whether a spectrogram of audio should be treated as a separate modality or simply as an image.

02What is multimodal learning?#

Multimodal learning is the branch of machine learning that builds models capable of processing and relating information from two or more modalities at once. Baltrušaitis, Ahuja, and Morency (2019) identify five core challenges in the field: representation (encoding different data types in a shared form), translation (converting one modality into another, such as image captioning), alignment (matching related parts across modalities, such as a spoken word to a video frame), fusion (combining modalities to make a prediction), and co-learning (using knowledge from one modality to help learn another). A landmark example is OpenAI's CLIP (Radford et al., 2021), which was trained on 400 million image–text pairs to place pictures and their descriptions close together in a shared embedding space, allowing it to recognize images from plain-language descriptions it was never explicitly trained on.

Uncertainty Factor: It is hard to know whether a multimodal model truly understands the connection between modalities or is relying on shortcuts, such as answering a question about an image using only the wording of the question while ignoring the picture itself.

03Where is multimodal learning used? Can you find one real application and identify the types of information it combines?#

Multimodal learning is used in autonomous vehicles, medical diagnosis (combining scans with patient records), virtual assistants, video captioning, accessibility tools for blind and low-vision users, and AI chat assistants that can read images.

Real application — the Waymo Driver (autonomous ride-hailing): Waymo's self-driving vehicles, operating as a public robotaxi service in cities such as Phoenix and San Francisco, fuse several sensor modalities to build a single picture of the road:

  • Cameras (vision): Read traffic lights, road signs, lane markings, and color information.
  • LiDAR (3D depth): Laser pulses produce a 3D point cloud measuring the exact shape and distance of cars, cyclists, and pedestrians, including at night.
  • Radar (velocity): Radio waves measure the speed and distance of objects and keep working in rain, fog, and dust.
  • Microphones (audio): External audio sensors detect sirens so the vehicle can yield to emergency vehicles.
  • Maps and positioning data: Detailed prior maps plus GPS and motion sensors locate the car within its lane.

Each sensor covers another's weakness: cameras struggle in darkness and glare, LiDAR cannot read a traffic light's color, and radar has low spatial resolution. Combining them produces a more reliable understanding than any single sensor could.

Uncertainty Factor: Waymo does not publish the full internal architecture of its fusion system, so the precise way and stage at which these sensor streams are combined inside the model cannot be confirmed from public information.

04What is embodied AI?#

Embodied AI refers to artificial intelligence that has a physical or simulated body and learns by perceiving and acting within an environment, rather than only analyzing static datasets. Instead of answering a question about a photo, an embodied agent must move through a room, pick up objects, and see the consequences of its actions. The idea draws on the hypothesis that intelligence emerges from interaction between an agent and its surroundings. Examples include household and warehouse robots, self-driving cars, drones, and virtual agents trained in 3D simulators such as AI2-THOR and Habitat before being transferred to real robots (Duan et al., 2022).

Uncertainty Factor: Skills learned in simulation often fail to transfer to the real world (the "sim-to-real gap") because real friction, lighting, and object textures differ from their simulated versions, making real-world performance difficult to predict in advance.

05How are multimodal learning and embodied AI connected?#

Embodied AI depends on multimodal learning because the physical world is inherently multimodal. A robot that must "pick up the red cup next to the sink" needs to understand the language of the instruction, see the scene through cameras, estimate depth to judge distance, and feel force through touch sensors to grip without crushing the cup. In return, embodiment gives multimodal models something static datasets cannot: grounding. By acting and observing outcomes, an agent learns what words like "heavy," "behind," or "fragile" physically mean. Google DeepMind's RT-2 (Brohan et al., 2023) illustrates this connection directly: it is a vision-language-action model that takes camera images and text instructions as input and outputs robot movement commands, transferring knowledge learned from web images and text into physical robot control.

Uncertainty Factor: It remains uncertain whether scaling up web-trained multimodal models is enough to produce reliable general-purpose robots, or whether robots will still need large amounts of costly, hands-on physical training data for every new task.

References

  • Baltrušaitis, T., Ahuja, C., & Morency, L.-P. (2019). Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(2), 423–443. https://doi.org/10.1109/TPAMI.2018.2798607
  • Brohan, A., et al. (2023). RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv:2307.15818. https://arxiv.org/abs/2307.15818
  • Duan, J., Yu, S., Tan, H. L., Zhu, H., & Tan, C. (2022). A survey of embodied AI: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence, 6(2), 230–244. https://doi.org/10.1109/TETCI.2022.3141105
  • Radford, A., et al. (2021). Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning (ICML), 8748–8763.
  • Waymo. The Waymo Driver: Technology overview. https://waymo.com/waymo-driver/
Back to all sets

Set 07 · AI & Robotics

AI & Autonomous Robots

From 70 years of AI to a robot guarding Mall of Qatar

5 questions8 min readProf. Gianni

01How do you define AI?#

Artificial intelligence (AI) is the branch of computer science that builds systems able to perform tasks that would normally require human intelligence: perceiving the world, reasoning, learning from experience, making decisions, and acting on them. The term was coined by John McCarthy and his colleagues in their 1955 proposal for the 1956 Dartmouth Summer Research Project, where McCarthy later described AI as "the science and engineering of making intelligent machines." The standard textbook by Russell and Norvig frames AI more precisely as the study of rational agents: systems that sense their environment and choose the actions expected to best achieve their goals.

My own working definition is: AI is software that turns data and goals into good decisions under uncertainty. Almost every AI system in use today is narrow AI, meaning it is very good at one task (translating text, recognizing faces, playing chess). A system that matches humans across every kind of task, often called artificial general intelligence (AGI), does not yet exist.

Uncertainty Factor: There is no single agreed definition of AI, and the boundary keeps moving. This is known as the "AI effect": once a problem such as chess or route planning is solved, people often stop calling the solution AI. Researchers also disagree about whether today's large language models truly "reason" or only imitate reasoning from statistical patterns.

02Can you name at least three different sub-fields of AI?#

  • Machine Learning (ML): Algorithms that learn patterns from data instead of following hand-written rules. Deep learning, which uses large neural networks, is the sub-area behind most recent breakthroughs. Example: spam filters and recommendation systems.
  • Computer Vision: Teaching machines to interpret images and video, including object detection, face recognition and scene understanding. Example: a self-driving car recognizing pedestrians and traffic lights.
  • Natural Language Processing (NLP): Understanding and generating human language. Example: machine translation, voice assistants and chatbots such as ChatGPT.
  • Robotics: Combining perception, planning and control so that machines can act in the physical world. Example: warehouse robots and surgical robots.
  • Knowledge Representation, Reasoning and Planning: Representing facts and rules so a system can draw conclusions and plan sequences of actions. Example: route planners and scheduling systems.
Uncertainty Factor: These sub-fields overlap more every year. A modern robot uses machine learning, computer vision and planning at the same time, and multimodal models now process text and images together, so the lines between sub-fields are becoming blurry.

03AI has been around for about 70 years so far. Why is it booming right now?#

AI has gone through cycles of hype and disappointment, including two "AI winters" (in the 1970s and the late 1980s) when funding dried up because the technology could not deliver on its promises. The current boom comes from several factors arriving at the same time:

  • Massive amounts of data: The internet, smartphones and sensors produce enormous datasets. ImageNet, released in 2009 with more than 14 million labeled images, gave researchers a shared benchmark for training vision models.
  • Cheap, powerful computing: Graphics processing units (GPUs), originally built for video games, turned out to be ideal for the parallel math inside neural networks. In 2012, AlexNet, a deep network trained on two GPUs, won the ImageNet competition with a top-5 error rate of 15.3%, compared with 26.2% for the runner-up. That result convinced the field to switch to deep learning. Cloud computing now lets anyone rent this power.
  • Better algorithms: Improvements in training deep networks, and especially the Transformer architecture introduced in 2017 ("Attention Is All You Need"), made it possible to train very large language models that power tools like ChatGPT.
  • Investment and easy access: When ChatGPT launched in November 2022, it reached an estimated 100 million users in about two months, one of the fastest-growing consumer applications ever. That public excitement brought huge investment from companies and governments, which funds even larger models.
Uncertainty Factor: It is unclear whether this boom will continue or lead to another AI winter. Training the largest models requires enormous amounts of energy, money and data, and some researchers argue that simply scaling up current methods will eventually stop producing large improvements.

04Can you name at least three application sectors where robots are being widely employed? What are the reasons?#

  • Manufacturing: This is the largest market for robots. According to the International Federation of Robotics, more than 4 million industrial robots were operating in factories worldwide in 2023, with the automotive and electronics industries as the biggest users. Reasons: robots weld, paint and assemble with very high precision, work 24/7 without fatigue, and take over dangerous jobs such as handling hot metal or toxic paint.
  • Logistics and warehousing: Fleets of mobile robots move shelves and packages inside fulfillment centers. Amazon reported more than 750,000 mobile robots working in its operations in 2023. Reasons: the explosive growth of e-commerce, the need for fast same-day delivery, labor shortages, and the fact that warehouses are structured indoor environments where robots can navigate reliably.
  • Healthcare: Surgical systems such as the da Vinci robot let surgeons perform minimally invasive operations through tiny incisions, and hospitals also use robots for delivering medicine, disinfection and rehabilitation. Reasons: greater precision and steadiness than the human hand, smaller incisions and faster patient recovery, and reduced exposure of staff to infection.

Agriculture is a fast-growing fourth sector. In 2024 my own team built an Arduino–Raspberry Pi automated seeder that planted 3× faster than manual seeding, which showed me how robots can help farms facing labor shortages and the need for more precise use of seeds, water and fertilizer.

Uncertainty Factor: Robot adoption is very uneven across countries and industries, because robots are expensive to buy, program and maintain. It is still uncertain how quickly robots will move from structured spaces like factories and warehouses into messy, unpredictable environments like homes, streets and farms.

05Can you identify three major challenges for a wheeled autonomous robot performing a 24h surveillance task in a large facility? (e.g., something like Mall of Qatar)#

Mall of Qatar is one of the largest malls in the country, with hundreds of stores spread over several floors, large open atriums, glass storefronts and crowds that change throughout the day and night. A wheeled security robot patrolling it around the clock would face three major challenges:

  • 1. Localization and navigation in a huge, crowded, changing space: GPS does not work reliably indoors, so the robot must build and use its own map (SLAM, simultaneous localization and mapping) from LiDAR, cameras and wheel odometry. Long, repetitive corridors look alike and can confuse the robot about where it is, glass storefronts and polished floors reflect or absorb laser beams, and store layouts and seasonal decorations change the map. During the day it must move safely through dense crowds, children and strollers, and because it has wheels it cannot climb stairs or use escalators, so it needs elevator access or one robot per floor.
  • 2. Energy and 24-hour endurance: Batteries cannot power motors, sensors and onboard computers for a full day, so the robot must recognize when its charge is low, return to a charging dock by itself and plug in without help. While it charges, parts of the mall are unguarded, so the patrol schedule (or a team of robots taking turns) must be planned to keep coverage continuous. Running nonstop also wears out wheels, motors and sensors, so the system needs monitoring and maintenance plans.
  • 3. Reliable perception and detecting real threats: The robot has to recognize genuinely unusual events (an intruder after closing time, a fire, smoke, water leaks, an unattended bag) while ignoring normal behavior. Lighting changes dramatically between bright daytime atriums and dim night-time corridors, and people and objects often block the camera's view. Too many false alarms make security staff ignore the robot; too few means real incidents are missed.

A fourth challenge cuts across all three: safety, privacy and trust. Real deployments have shown what can go wrong. In 2016 a Knightscope K5 security robot collided with a toddler at a shopping center in California, and in 2017 another K5 rolled into an office-complex fountain in Washington, D.C. Constantly recording shoppers also raises privacy questions, which in Qatar fall under Law No. 13 of 2016 on protecting personal data.

Uncertainty Factor: It is uncertain whether a robot can be trusted to make security decisions on its own. Most current systems keep a human operator in the loop to confirm alarms, and the right balance between autonomy, cost, and human oversight for a facility as large as Mall of Qatar has not been settled.

References

  • McCarthy, J., Minsky, M. L., Rochester, N., & Shannon, C. E. (1955). A proposal for the Dartmouth Summer Research Project on Artificial Intelligence.
  • Russell, S., & Norvig, P. (2021). Artificial Intelligence: A Modern Approach (4th ed.). Pearson.
  • Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., & Fei-Fei, L. (2009). ImageNet: A large-scale hierarchical image database. IEEE Conference on Computer Vision and Pattern Recognition (CVPR).
  • Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2012). ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems 25.
  • Vaswani, A., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems 30.
  • Hu, K. (2023, February 2). ChatGPT sets record for fastest-growing user base – analyst note. Reuters.
  • International Federation of Robotics. (2024). World Robotics 2024 – Industrial Robots.
  • Thrun, S., Burgard, W., & Fox, D. (2005). Probabilistic Robotics. MIT Press.
Back to all sets