Tip of the Day

Do not go where the path may lead, go instead where there is no path and leave a trail.

Showing posts with label Machine Learning. Show all posts
Showing posts with label Machine Learning. Show all posts

Introduction to Machine Learning: Definition, Types, and Applications

Introduction to Machine Learning: Definition, Types, and Applications


 

Introduction

Machine learning, a subfield of artificial intelligence (AI), is the driving force behind many modern technologies, from chatbots and predictive text to autonomous vehicles and medical diagnostics. It's a field that enables computers to learn without being explicitly programmed, and it's changing every industry. This article explores the definition, types, and applications of machine learning, providing insights into its potential and limitations.


What is Machine Learning?

Machine learning is defined as the capability of a machine to imitate intelligent human behavior. It's a way to use AI, allowing computers to recognize visual scenes, understand natural language, or perform actions in the physical world. Arthur Samuel, an AI pioneer, defined machine learning in the 1950s as "the field of study that gives computers the ability to learn without explicitly being programmed."


Types of Machine Learning

Machine learning can be categorized into three main types:


Supervised Machine Learning

 Models are trained with labeled data sets, allowing them to learn and grow more accurate over time. For example, an algorithm trained with pictures of dogs and other objects can identify pictures of dogs on its own.


Unsupervised Machine Learning

 This type looks for patterns in unlabeled data, finding trends that people aren't explicitly looking for. It can analyze online sales data to identify different types of clients, for example.


Reinforcement Machine Learning

Machines are trained through trial and error using a reward system. It can train models to play games or drive autonomous vehicles by reinforcing correct decisions.

Types of Machine Learning




Applications of Machine Learning

Machine learning has a wide range of applications across various sectors:


Business: From manufacturing to retail, machine learning unlocks new value and boosts efficiency. 67% of companies are using machine learning, according to a recent survey.

Healthcare: Machines can diagnose medical conditions based on images, offering insights that may be beyond human capability.

Environment: Concerns about the economic and environmental sustainability of deep learning, a subset of machine learning, are being addressed to ensure responsible usage.

Entertainment: Platforms like Netflix use machine learning for personalized suggestions, enhancing user experience.


Ethical Considerations

Machine learning also brings social, societal, and ethical implications. It's vital to engage with these tools responsibly, considering how to use them for the good of all. Understanding the potential and limitations of machine learning is essential for leaders across industries.


Conclusion

Machine learning is not just a technological advancement; it's a paradigm shift that's influencing every aspect of our lives. From understanding its basic principles to recognizing its potential and limitations, machine learning is a field that no one can afford to ignore. Its applications are vast, and its impact is profound, shaping the future of work, healthcare, entertainment, and more.

Unlocking the Power of Mathematics in Machine Learning

Unlocking the Power of Mathematics in Machine Learning

 Machine learning has rapidly gained popularity in recent years, as businesses and organizations look for innovative ways to solve complex problems and make better predictions. But behind every successful machine learning model is a foundation of mathematical concepts. Linear algebra, calculus, probability, and statistics are just a few of the mathematical disciplines that play a crucial role in the development and implementation of machine learning algorithms.

If you're interested in exploring the mathematical foundations of machine learning, there are many courses available online, both paid and free. Here, we'll take a look at some of the best options for unlocking the power of mathematics in machine learning.


The top 5 popular free courses for Mathematics for Machine Learning include:

  1.  Mathematics for Machine Learning by Coursera : This course is a great starting point for anyone looking to understand the mathematical concepts that are fundamental to machine learning. With a focus on problem-solving and critical thinking, this course is perfect for those who are new to mathematics or looking to brush up on their skills.
  2. Linear Algebra - Foundations to Frontiers by edX: https://www.edx.org/learn/linear-algebra
  3. Introduction to Linear Algebra by MIT OpenCourseWare: https://ocw.mit.edu/courses/mathematics/18-06-linear-algebra-spring-2010/
  4. Machine Learning by University of Washington on Coursera: https://www.coursera.org/courses?query=machine%20learning%20university%20of%20washington
  5. Mathematics for AI by IBM on Coursera: https://www.coursera.org/courses?query=mathematics%20for%20AI%20IBM

Note that some courses may have limited content available for free, but still provide a good introduction to the topic.


The top 5 popular Paid courses for Mathematics for Machine Learning include:

1. Mathematics for Machine Learning by Coursera: https://www.coursera.org/courses?query=mathematics%20for%20machine%20learning

This comprehensive course covers the mathematical concepts that are essential for machine learning, including linear algebra, calculus, and probability. With practical examples and hands-on exercises, this course is perfect for those looking to deepen their understanding of the math behind machine learning.

2. Linear Algebra and Learning from Data by edX: https://www.edx.org/learn/linear-algebra-and-learning-from-data

This course focuses specifically on linear algebra and its role in machine learning. With a blend of theory and practical examples, this course is perfect for anyone looking to understand the underlying mathematics behind machine learning models.

3. Deep Learning Specialization by Coursera: https://www.coursera.org/specializations/deep-learning

 This specialization covers all aspects of deep learning, including the mathematical concepts that are essential for understanding and building neural networks. From linear algebra to backpropagation, this course is perfect for anyone interested in taking a deep dive into the math behind deep learning.

4. Introduction to Mathematical Thinking by Stanford on Coursera: https://www.coursera.org/courses?query=introduction%20to%20mathematical%20thinking%20stanford

Introduction to Mathematical Thinking by Stanford on Coursera is a comprehensive course designed for anyone looking to gain a deeper understanding of mathematical concepts and problem-solving techniques. This course covers a wide range of mathematical topics, from basic set theory to more advanced concepts such as graph theory and linear algebra. Throughout the course, students will engage in interactive exercises and problem-solving activities that will help them develop their critical thinking and problem-solving skills. Whether you're a beginner looking to build a strong foundation in mathematics or an experienced practitioner looking to brush up on your skills, this course is the perfect starting point for unlocking the power of mathematical thinking. With engaging video lectures, hands-on exercises, and a supportive community of learners, 

5. Multivariate Calculus & Deep Learning by fast.ai: https://course.fast.ai/ml.html

This course focuses specifically on the role of multivariate calculus in deep learning. With hands-on examples and practical exercises, this course is perfect for anyone looking to understand the math behind neural networks and other deep learning models.

It is best to choose a course based on your current mathematical knowledge and familiarity with programming.


Here are 10 popular YouTube video courses for Mathematics for Machine Learning:


1. 3Blue1Brown's Essense of Linear Algebra

2. Neural Networks Demystified by Lazy Programmer Inc

3. Machine Learning with Phil by Phil Tabor

4. Siraj Raval's Math of Intelligence

5. Deep Learning 101 by deeplizard

6. Fast AI's Practical Deep Learning For Coders

7. Machine Learning Mastery by Jason Brownlee

8. AI Adventures by Google Developers

9. Deep Learning Wizard by The Lazy Programmer

10. Applied AI Course by Kirill Eremenko

These video courses cover various mathematical concepts including linear algebra, calculus, and probability, and their applications in machine learning.








Shall We Let Computers Measure Beauty?

Shall We Let Computers Measure Beauty?

 As we all know, tastes differ and change over time. However, each epoch tried to define its own criteria for beauty and aesthetics. As science was developing, so was the urge to measure beauty quantitatively. Not surprisingly, the recent advancements in Artificial Intelligence pushed forward the question of whether intelligent models can overcome what seems to be human subjectivity.

A separate subfield of artificial intelligence (AI), called ‘computational aesthetics’, was created to assess beauty in domains of human creative expression such as music, visual art, poetry, and chess problems. Typically, it uses mathematical formulas that represent aesthetic features or principles in conjunction with specialized algorithms and statistical techniques to provide numerical aesthetic assessments. Computational aesthetics merges the study of art appreciation with analytic and synthetic properties to bring into view the computational thinking artistic outcome.

Brief History of Computational Aesthetics

Though we are used to thinking about Artificial Intelligence as a recent development, computational aesthetics can be traced back as far as 1933, when American mathematician George David Birkhoff in “Aesthetic Measure” proposed the formula M = O/C where M is the “aesthetic measure,” O is order, and C is complexity. This implies that orderly and simple objects appear to be more beautiful than chaotic and/or complex objects. Order and complexity are often regarded as two opposite aspects, thus, order plays a positive role in aesthetics while complexity often plays a negative role. Birkhoff applied that formula to polygons and artworks as different as vases and poetry, and is considered to be the forefather of modern computational aesthetics.

In the 1950s, German philosopher Max Bense and French engineer Abraham Moles independently combined Birkhoff’s work with Claude Shannon’s information theory to develop a scientific means of grasping aesthetics. These ideas found their niche in the first computer-generated art but did not feel close to human perception.

In the early 1990s, the International Society for Mathematical and Computational Aesthetics (IS-MCA) was founded. This organization is specialized in design with an emphasis on functionality and aesthetics and attempts to be a bridge between science and art.

In the 21st century, computational aesthetics is an established field with its own specialized conferences, workshops, and special issues of journals uniting researchers from diverse backgrounds, particularly AI and computer graphics.

Objectives of Computational Aesthetics

The ultimate goal of computational aesthetics is to develop fully independent systems that have or exceed the same aesthetic “sensitivity” and objectivity as human experts. Ideally, machine assessments should correlate with human experts’ assessment and even go beyond it, overcoming human biases and personal preferences.

Additionally, those systems should be able to explain their evaluations, inspire humans with new ideas, and generate new art that could lie beyond typical human imagination.

Finally, computing aesthetics can also provide a deeper understanding of our aesthetic perception.

In practical terms, computational aesthetics can be applied in various fields and for various purposes. To name a few, aesthetics can be used in the following applications:

  • as one of the ranking criteria for image retrieval systems;
  • in image enhancement systems;
  • managing image or music collections;
  • improving the quality of amateur art;
  • distinguishing between videos shot by professionals and by amateurs;
  • aiding human judges to avoid controversies, etc.

Features

The backbone of all classifiers is a robust selection of features that can be associated with the perception of a certain form of art. In the search for correlation with human perception, aesthetic systems apply specific sets of features for visual art and music that are developed by theorists in arts and domain experts.

Visual Art

Image aesthetic features could be categorized as low-level or high-level plus composition-based. However, some research is based on features related to saliency (Zhang and Sclaroff, 2013), object (Roy et al., 2018), and information theory (Rigau,‎1998). The selection of features largely depends on the type of art and the level of abstraction, as well as the algorithm applied. For instance, photography assessment relies heavily on the compositional aspects, while measurement of the beauty of abstract art requires another approach assessing color harmony or symmetry (Nishiyama et al.,2011).

Low-level features try to describe an image objectively and intuitively with relatively low time and space complexity. They include color, luminance and exposure, contrast, intensity, edges, and sharpness.

High-level features include regions and contents as aspects that make great contributions to overall human aesthetic judgment and try to establish the regions of an image that seem to be more important for human judgment and find the correlation between the content and human reaction.

Composition-based features differ for photography and artwork and may include depending on the form of art a range of features, such as Rules of Thirds, Golden Ratio (Visual Weight Balance), focus and focal length, ISO speed rating, geometric composition and shutter speed (Aber et al., 2010).

Music

Similarly to image analysis, music aesthetics assessments try to combine research in

human perception and cognition of basic dimensions of sound, such as loudness or pitch and in higher-level concepts related to music, including the perception of its emotive content (Juslin and Laukka, 2004), as well as performance specific traits (Palmer, 1997) to develop a comprehensive set of features that would be able to assess a piece of music.

In 2008, Gouyon et al. offered a hierarchy organized in three levels of abstraction starting from the most fundamental acoustic features, to be extracted directly from the signal, and progressively building on top of them to get to model more complex concepts derived from music theory and even from cognitive and social phenomena:

Low-level features are related to the physical aspect of the signal and include loudness, pitch, timbre, onsets, and rhythm (e.g., see Justus and Bharucha, 2002).

Mid-level features move to a higher level of abstraction within the music theory and cover tempo, tonality, modality, etc.

High-level features try to establish a correlation between abstract music descriptors like genre, mood, and instrumentation and human perception.

Methods and Algorithms

At its broadest, we can speak of computational aesthetics as a tool to assess aesthetics in visual art or music and as a means to generate new art.

For aesthetics assessment, various algorithms have been proposed over the past few years based either on classification or clusterization.

Classification approach

There are a number of algorithms that are extensively used to assess image aesthetics by means of classification. Among the most popular are AdaBoost, Naive Bayes, and Support Vector Machine, and substantial work is also conducted using Random Forests and Artificial Neural Networks (ANNs).

AdaBoost in computational aesthetics is a widely used method that is believed to render the best results. It was first offered in 2008 by Luo and Tang who conducted a study on photo quality evaluation, with the unique characteristic of focusing on the subject. They utilized Gentle AdaBoost (Torralba et al., 2004), a variant of AdaBoost that uses a specific way of weighting its data, applying less weight to outliers. The success rate obtained was 96%. However, when Khan and Vogel (2012) utilized their proposed set of features for photographic portraiture aesthetic classification, the accuracy rate with the multiboosting variant (multi-class version) of AdaBoost fell to 59.14% (Benbouzid et al., 2012).

Naïve Bayes is another popular method that was used in the same study by Luo and Tang (2008). In 2009, Li and Chen utilized the Naïve Bayes classifier to aesthetically classify paintings in which the results were described as robust. The success rate achieved utilizing a Bayesian classifier was 94%.

Support Vector Machine is probably the most wide-spread algorithm for binary classification in computational aesthetics. It has been used since 2006 when Datta et al. studied the correlation between a defined set of features and their aesthetic value, by using a previously rated set of photographs and showed up to 76% of accuracy. Other studies that rested on the same classifier include Li and Chen (2009) who aesthetically classified paintings; Wong and Low (2009) who built a classification system of professional photos and snapshots, Nishiyama et al. (2011) who conducted a research on the aesthetic classification of photographs based on color harmony, and others, with an average accuracy rate of about 75% and higher.

Random Forest, though usually showing lower results as compared to Bayesian classifiers or AdaBoost, were used in a number of studies of photograph aesthetics. For instance, Ciesielski et al. (2013) achieved a 73% accuracy to assess photograph aesthetics. Khan and Vogel (2012) utilizing their proposed set of features for photographic portraiture aesthetic classification, achieved an accuracy of 59.79% by making use of random forests (Breiman, 2001).

Artificial Neural Networks (ANNs) rendered extremely good results when used with compression-based features by Machado et al. (2007) and Romero et al. (2012). The former research aimed at the identification of the author of a set of paintings and reported a success rate from 90.9% to 96.7%. The latter work used an ANN classifier to predict the aesthetic merit of photographs at a success rate of 73.27%.

Convolutional Neural Networks (CNNs) are state-of-the-art deep learning models for rating image aesthetics that have been extensively used in the past few years. CNNs learn a hierarchy of filters, which are applied to an input image in order to extract meaningful information from the input. For example, Denzler et al. (2016) applied the AlexNet model (Krizhevsky et al., 2012) on different datasets to experimentally evaluate how well pre-learned features of different layers are suited to distinguish art from non-art images using an SVM classifier. They report the highest discriminatory power with a Network trained on the ImageNet dataset, which outperforms a network solely trained on natural scenes.

Clustering

Image clustering is a very popular unsupervised learning technique. By grouping sets of image data in a particular way, it maximizes the similarity within a cluster, simultaneously minimizing the similarity between clusters. In computational aesthetics, researchers use K-Means, Fuzzy Clustering, and Spectral Clustering in image analysis.

K-Means Clustering is widely used to analyze the color scheme of an image. For instance, Datta et al. (2006) used k-means to compute two features to measure the number of distinct color blobs and disconnected large regions in a photograph. Lo et al. (2012) utilized this method to find dominant colors in an image.

Fuzzy Clustering is a form of clustering in which each data point can belong to more than one cluster, therefore it is used in multi-class classification (see, for example, Felci Rajam and Valli (2011)). Celia and Felci Rajam (2012) utilized FCM clustering for effective image categorization and retrieval.

Spectral Clustering is used to identify communities of nodes in a graph based on the edges connecting them. In computational aesthetics, a spectral clustering technique named normalized cuts (Ncut) was used to organize images with similar feature values (Zakariya et al., 2010).

Generative models

A separate task of computational aesthetics is to generate artwork independently from human experts. At present, the algorithm that is best known for directly learning the transformations between images from the training data is Generative Adversarial Network(GAN). GANs automatically learn the appropriate operations from the training data and, therefore, have been widely adopted for many image-enhancement applications, such as image super-resolution and image denoising. Machado et al. (2015) also used GANs for automatically enhancing image aesthetics by performing mainly tone adjustment.

Example that combines the content of a photo with a well-known artwork

Conclusion: Restrictions and Limitations

Aspiring to reach objectivity, research in computational aesthetics tries to reduce the focus to form, rather than to content and its associations to a person’s mind and memories. However, from a psychophysiological viewpoint, it is not clear whether we can have a dichotomy here or whether aesthetics is intrinsically subjective.

Besides, it is difficult to ascertain whether a system that performs on the same level as a human expert is actually using similar mechanisms as the human brain and, therefore, whether it reveals something about human intelligence.

It might be that in the future we will rely on machines in our artistic preferences, but for now, human experts will dictate their opinions and try to get machines simulate their choices.

Source: https://medium.com/sciforce/computational-aesthetics-shall-we-let-computers-measure-beauty-db2205989fb

What Machine Learning skills should I be learning now to set myself up for success in the coming years?

What Machine Learning skills should I be learning now to set myself up for success in the coming years?
Strong understating of the fundamentals - the ML concepts and algorithms, and the underlying math:
  • How Forward feed and backwards prop work.
  • The various loss functions and their considerations
  • The various activation functions and why they are needed
  • Optimization functions and why they are needed
  • Bias and variance / over and under fitting - what causes them, and the various methods to handle them
  • CNNs, RNNs, GANs, attention, Transformer, unsupervised and semi supervised, RL, decision trees, Ensemble Learning, SVM, Auto encoders…
  • Understand interpretation, bias, fairness
  • The statistics theory (the more the better), and the linear algebra and calculus technicalities
I highly recommend the "Neural Networks For Machine Leaning" course from University of Toronto, given by Geoffrey Hinton. It's a bit out dated in some not-so-meaningful sense, and definitely much harder than any other ML course out there. But if you survive through it, it provides deep mathematical intuition into ML, like no other course does.
It's a lot and not very easy, but if you do it - it will pay off. The libraries, frameworks, and hopefully also the concepts and algorithms will change over time. But if you have a solid understanding of the above, it will be very easy for you to keep up with the developments, grow, and adapt.

What is Google's capsule network? How is it different than convolutional neural networks?

What is Google's capsule network? How is it different than convolutional neural networks?
Capsules introduce a new building block that can be used in deep learning to better model hierarchical relationships inside of internal knowledge representation of a neural network. Intuition behind them is very simple and elegant.
Hinton and his team proposed a way to train such a network made up of capsules and successfully trained it on a simple data set, achieving state-of-the-art performance. This is very encouraging.
Nonetheless, there are challenges. Current implementations are much slower than other modern deep learning models. Time will show if capsule networks can be trained quickly and efficiently. In addition, we need to see if they work well on more difficult data sets and in different domains.
In any case, the capsule network is a very interesting and already working model which will definitely get more developed over time and contribute to further expansion of deep learning application domain. 
https://arxiv.org/pdf/1710.09829... (paper from Hinton et al. proposing capsule networks in 2017).



The main point is that while CNNs are great in recognizing both simple features in the lower layers of an image (edges and colors gradients), as well as their complex compositions in the deeper levels (ball, dog, cat, face, wheel, car …) - they do a very poor job in representing their rotational and translational relationships (meaning, how they are organized in space in respect to each other).
For example, a CNN might be very successful in recognizing the different elements of a face in an image - eyes, nose, mouth, and deduce that an image segment containing these - most probably represents a face. However, it will not be sensitive to the arrangement of the entities (mouth under nose, then two symmetric eyes above that), and might mistakenly recognize different arrangements of these entities also as face.
Hinton is especially critical of the mechanism that CNNs use to handle some translational invariance - the MaxPooling.
There are multiple consequences for this shortcoming, but two major ones are:
  • Misclassification of images that contain the “right” entities in a wrong pose, as explained above. Moreover, knowing that all the entities are arranged in a very specific relationship to one another (mouth under nose under eyes) - is a much stronger signal for the existence of a face, compared to just knowing they are there.
  • Inefficient representation that leads to ineffective learning - instead of having a small canonical set per entity + pose information, every pose of the entity is modeled separately. That leads to a huge training set, which is orders of magnitude larger than what’s required for a human brain to learn the same classification / recognition.
Note : Write your oprnion in comment box to improve the post and also to get better understanding to other reader .

Is strong AI inevitable?

Is strong AI inevitable?


Any question about the future is susceptible to unknown unknowns, futile speculations about things undiscovered. We need examples we can observe and interrogate now. While we don’t have strong AI, we do have rigorous examples of weak and strong intelligence. A comparison of the nature of knowledge in its weak and strong forms offers a penetrating and non-technical look into the prospects for strong AI.
Our last stop on this tour of the AI landscape introduced induction as the prevailing theory of knowledge creation, and the central role that explanations play in workable inductive systems. Here, I’ll apply that framework to shed some light on one of the most contentious and important debates in AI: Are we on the path to artificial general intelligence? Is tomorrow’s strong AI the inevitable extension of today’s weaker examples?
Here’s the plan: We’ll examine two points along the knowledge hierarchy, one associated with weak AI, the other a much stronger form. I’ve labelled these points predictions and explanations, respectively, and I’ll make these terms more precise as we go. Through concrete examples, you can evaluate the quality of each intelligence yourself, and decide whether the path from weak to strong seems smooth and incremental, or perilous and disjoint.

Possible is not inevitable

Before we dive in, there’s one aspect of this debate we should set aside, the difference between possible and inevitable. There’s a reasonable expectation that strong AI is physically possible. The physicist and pioneer of quantum computing David Deutsch provides a rich discussion of that possibility in The Fabric of Reality. The claim to possibility is grounded in the Turing principle: “There exists an abstract universal computer whose repertoire includes any computation that any physically possible object can perform.” Life embodies intelligence. The Turing principle says that a computer could be tractably built and programmed to render any physical embodiment, including objects such as intelligence-producing brains. The possibility of strong AI follows logically from the Turing principle, which Deutsch maintains is so widely accepted as to be pragmatically true.
“The Turing principle guarantees that a computer can do everything a brain can do. That is of course true, but it is an answer in terms of prediction, and the problem is one of explanation.” David Deutsch
Despite this principled argument, the inevitability of AI remains a hotly contested topic among philosophers, scientists and mathematicians (Roger Penrose being a particularly formidable critic). However, the disagreements are rooted less in theoretical arguments as in the specific “how do we get there from here” implications. Deutsch explains, “It is then not good enough for artificial-intelligence enthusiasts to respond brusquely that the Turing principle guarantees that a computer can do everything a brain can do. That is of course true, but it is an answer in terms of prediction, and the problem is one of explanation. There is an explanatory gap.”
To assess the inevitability of AI, we need to illuminate the explanatory gaps between weak and strong forms of intelligence.

When everything is prediction

One last bit of housekeeping. For some, the inevitability of AI is not only an answer in terms of prediction, prediction is the essence of intelligence itself. This is patently untrue. Intelligence is like Whitman, “I am large, I contain multitudes.” But it’s the law of the hammer to treat everything as if it were a nail. So motivated, any functional gap in AI may be framed as a prediction problem: The crux of the AI problem is prediction under uncertainty; reasoning entails activities of prediction; predicting general rules as common sense; and so on.
When everything is rooted in prediction, including the goal, every open problem appears incremental. Frequently, arguments for inevitability reference the indirect factors of production in prediction engines, such as the pace of investment and increases in people, computing resources and data. This is the tweet-sized version of the AI roadmap: Prediction is a unit of intelligence. To achieve greater intelligence, just add more prediction resources.
In a recent article for Harvard Business Review, one of the most influential AI researchers Andrew Ng offers this rule of thumb: “If a typical person can do a mental task with less than one second of thought, we can probably automate it using AI either now or in the near future.” Data and the talent to expertly apply the software are the only scarce resources impeding progress for these one-second tasks. To be fair, Ng cuts through the hype, acknowledging that “there may well be a breakthrough that makes higher levels of intelligence possible, but there is still no clear path yet to this goal.”
The tendency to conflate lower order capabilities like predictions with more sophisticated forms of intelligence is not limited to sloppy talk in business and marketing forums. It extends deep into the technical domains of statistics and machine learning.
The statistician Galit Shmueli explains how the conflation of prediction and explanation has reached epidemic proportions, due to the indiscriminate use of statistical modeling. “While this distinction has been recognized in the philosophy of science, the statistical literature lacks a thorough discussion of the many differences that arise in the process of modeling for an explanatory versus a predictive goal.” For many, explanation has come to mean “efficient” or “interpretable” prediction.
Consider this example from the University of Washington, explaining the predictions of any classifier. “By ‘explaining a prediction’, we mean presenting textual or visual artifacts that provide qualitative understanding of the relationship between the instance’s components (e.g. words in text, patches in an image) and the model’s prediction.” The “explanation” is a human interpretable description derived from the more complex model.
LIME: Local Interpretable Model-agnostic Explanations. Here, explanation refers to a low level description (dashed black line) derived from a complex model (blue/pink background).
It’s an elegant solution for making complex models more interpretable, but as we’ll see, it’s a far cry from explanations in the more rigorous scientific sense of the term. Tellingly, Ng and his collaborators leveraged this approach for “explaining predictions” in a recent paper.
We don’t care about wordplay, we want to understand the substance of the thing. And whatever we call it, we need to determine whether something substantial separates the products of strong and weak intelligence. To do that, we’ll look in more detail at predictions, as associated with machine learning, and explanations, as associated with scientific discovery.
Machine learning is associated with predicting, the process of scientific discovery with explaining.

Weak Intelligence: Prediction engines

This is my dog Thor. His sense of smell is orders of magnitude better than mine. He hears sounds at frequencies 20,000 Hz above my upper limit. Thor perceives the world in a way that I could never apprehend. He builds remarkably versatile predictive models. He learns through positive and negative reinforcement (liver treats and time-outs, respectively). He frequently surprises me, when I can’t hear what he hears or smell what he smells. Yet never, ever, explains. When Thor’s world changes, he doesn’t ask why.
My dog Thor
Despite our inability to have deep conversations, I’ve come to love Thor. And I feel the same way about my other go-to prediction engine, the navigation app in my phone. It frequently delights me, routing me through traffic in ways I’ve never travelled before. Like Thor, the app is powered by a torrent of data beyond my senses, rapidly building and updating its directions. As explained on Google’s AI blog, “In order to provide the best experience for our users, this information has to constantly mirror an ever-changing world.” But it never offers any higher order explanation of traffic that reaches beyond its senses.
Proponents of prediction may cry foul. “These are your examples of prediction, a navigation app and your dog?!” Again, keep in mind that I’m only trying to locate a plateau of weak intelligence, such that we can contrast it with the stronger forms that follow. In that spirit, we need to set a boundary.
Peeking under the hood, the nature of predictions is revealed in their boundariesBrett Hall offers a simple illustration, a pot of boiling water. The temperature of the water increases steadily until the boiling point. Based on observations before the boiling point, it would be quite reasonable to predict the temperature will continue to rise. But once the temperature exceeds the boiling point, it defies prediction and demands explanation. Hall wryly asks, if extrapolation fails even when applied to this simple system, how can it be expected to succeed when things are more complicated?
You may be thinking, this is a toy model applied to a complex system. It’s bound to fail. Just give it more data and a more robust model! This is a fair criticism, as the entire premise of the incremental roadmap is that we can continue to ingest new data to create evermore sophisticated systems.
So let’s adopt the most idealized concept of a prediction engine we can imagine, an oracle with godlike powers of divination. You can ask the oracle whatever you want and it will predict what will happen. Of course, cash being king, you ask it to predict stock prices. And it works! You invest small amounts of money in stocks predicted to rise, and your investments increase. (Curiously, you notice a slight discrepancy between the predicted values and the actual stock prices, but you think nothing of it at the time.)
Gradually, the size and pace of your investments increase, as does your influence and reputation. Now, you’re not only playing the markets, you’re moving them! But strangely, your oracle begins to fail you. Those discrepancies between the predicted and actual prices are now quite pronounced, frequently undermining your investments. Your oracle’s predictions have degraded to approximations. It can still describe the system, but it can’t tell you the impact of your interventions on that system, without moving you inside the model.
These features encapsulate what the political scientist Eugene Meehancalled the system paradigm of explanation. He described explanations as empirical generalizations, formalized in models such as Bayesian networks and structural equations. Predicted outcomes are insufficient. Observations cannot simply fit the model, for a model can be created to fit any set of facts or data. If the system sufficiently reflects the environment, the prediction applies; if it doesn’t, the prediction fails. This is what it means to “mirror an ever-changing world.” And this problem is endemic to the task of prediction.
Implicitly, the system paradigm of explanation is what many people associate with prediction engines and the incremental path to improve them. However, as an exemplar of strong AI, we can do better.

Strong Intelligence: Scientific explanations

I’ve argued previously that science, our most successful knowledge creating institution, is an exemplar for artificial intelligence. With that in mind, I’ll surface examples that illustrate the scientific conception of explanations. As compared with prediction, these examples encompass a range of important differences, in form, function, reach and integration. I’ll also use one of the best scientific explanations, quantum mechanics, to illuminate these differences. If AI begins automating this quality of scientific discovery, we’ll all most certainly agree it’s strong indeed.
But fittingly, let’s start with a sunrise. For a very long time, we believed the sun rises because we observe it. But now we understand it in terms of the functioning of the solar system and the laws of physics. Explaining is much more robust than predicting. We know the sun is rising even when it’s cloudy. If we were orbiting the planet, frequent sunrises would not be at all surprising. (They would, however, remain awe-inspiring!) Appearances notwithstanding, we know in reality that the sun isn’t rising at all.
Let’s pause for a moment to behold this explained sunrise. There’s a profound distance between our observations of the sunrise and its explanation. The observations appear regular and uniform. Every morning the sun rises in the east. Yet the deeper explanation of the sunrise permits unobservable data, such as what’s happening when the sun is obscured by the clouds. It even admits the imagination, the counterfactual case of observers in orbit.
In his discussion of the complexity of scientific inference, the philosopher Wesley Salmon used this example to characterize predictions based on “crude induction” as “unquestionably prescientific”, even the antithesis of scientific explanations. Explanations sit at the apex of the knowledge hierarchy due to their depth and reach. “A scientific theory that merely summarized what had already been observed would not deserve to be called a theory.” Here in the scientific milieu, the illusion of induction is laid bare. Contrary to the idea that knowledge is induced from data, science reveals a rich integration of explanations that predicts otherwise unobservable data.
Let’s look under the hood of explanations, as we did with predictions. Explanations (or theories) consist of interpretations of how the world works and why. These explanations are expressed in formalisms as mathematical or logical models. Models provide the foundation for predictions, which in turn provide the means for testing through controlled experiments.
Adapted from David Deutsch, Apart from Universes
In this schema, explanations are the unity of interpretations, formalisms and predictions. Each component serves a functional role and each may stand-alone. Iteratively, explanations may be tested via their predictions, and their results may draw attention to explanatory gaps in need of attention. But their unifying purpose is to elucidate and criticize the explanations. In this light, predictions (a part) is subsumed by explanations (the whole). Explanations behave more like living ecosystems than static artifacts, exhibiting a dynamic churn of conjecture, predictions, experimentation, and criticism.
The complex nature of scientific explanations finds rich expression in quantum mechanics. There’s a risk that this example is more complicated than the thing I’m trying to explain. But this is also one of the most robust and counter-intuitive explanations in science. It truly captures the nature of alien-like superintelligence, our expectation of strong AI. Even grossly simplified, it serves our purpose here, as a peak of strong intelligence.
The formalisms of quantum mechanics may be used for prediction without any interpretation of the underlying physical reality. Yet there’s a wide range of competing interpretations, such as collapsing waves, pilot waves, many worlds, or many minds, to name but a few. And critically, these interpretations matter. They inspire researchers, influence research programs and contribute to the production of future knowledge.
Going a little deeper, consider the distinction between laws that precisely describe observations (phenomena), and explanations that have the reach and power to create those laws. The philosopher Nancy Cartwright, in her influential book, How the Laws of Physics Lie, highlights the difference between a generalized account capable of subsuming many observations (phenomenological laws), and the specificity needed for models to actually predict the real world. “The route from theory to reality is from theory to model, and then from model to phenomenological law. The phenomenological laws are indeed true of the objects in reality — or might be; but the fundamental laws are true only of objects in the model.”
So again, to appreciate how profoundly different scientific explanations are from predictions, let’s return to the example of quantum mechanics. There’s a generalized theory of how a system will change (formalized in Schrödinger’s equation) and the specific energies acting on the system (formalized as the Hamiltonian operator set up for the system). Cartwright explains, “It is true that the Schrödinger equation tells how a quantum system evolves subject to the Hamiltonian; but to do quantum mechanics, one has to know how to pick the Hamiltonian.”
Just like our sunrise, quantum theory spans a rich hierarchy of explanations, not data, extending from generalized concepts to more specific models capable of actually predicting outcomes.
Inductive systems deal with structure, too, such as the hierarchy of features that compose an image. But here, we’re talking about structures that don’t just describe or represent data, but rather explain their underlying causes. These integrations form a lattice that supports the whole edifice, extending all the way down to language itself as primitive explanations. Just as reality exists across compounding tiers of abstraction, explanations need the support of other explanations. And frequently, explanations stand quite removed from data and observations. Like some mind-body duality, this unity is the content of knowledge itself.

Is strong AI inevitable?

We now have two distinct vantage points, the weaker plateau of predictions and the stronger peak of scientific explanations. So off we go from one to the other, on our trek to answer our question.
The first impediment on our path is the gap of interventions. To predict is to infer new observations beyond what we’ve observed thus far. Predictions are bounded by an assumption of uniformity. If the system changes, new data are needed to build a more accurate model. A model is less valuable if it’s merely descriptive. We want predictions we can act on, to buy a stock or treat a disease. But intervening entails changing the system, breaking the assumption of uniformity.
“God asked for the facts, and they replied with explanations.” Judea Pearl
So how wide is the gap of interventions? Judea Pearl is a pioneer of modern AI and probabilistic reasoning. In The Book of Why, he places associative predictions on the lowest rung of intelligence; interventions and more imaginative counterfactual reasoning as stronger forms. Pearl explains why this higher order knowledge cannot be created from the probabilistic associations that characterize inductive systems. As one of the originators of the idea, he’s now “embarrassed” by the claim that probabilities comprise a language of causation and higher order knowledge. Suffice it say the gap of interventions is at best spanned by a rickety bridge that not even Indiana Jones would cross without pause. But we troddle on.
Eventually we arrive at even more imposing barrier, the cliff of data. While interventions pose a significant gap along our path, they’re nothing compared to this impediment. The fuel of inductive systems is data. Scientific conjectures are often described as leaps of imagination. The “data” of imagination, counterfactuals, are by definition not observed facts! Quite unlike data, knowledge is composed of a rich lattice of mutually supporting explanations. When I try to reconcile the enthusiasm for data-driven technologies with the reality of conjectural knowledge, I’m reminded of Wile E. Coyote running off the cliff. What’s holding him up?
In his critical appraisal of deep learning, the scientist and AI researcher Gary Marcus claims deep learning may be hitting a wall (or perhaps falling off the cliff). He surveys many of the topics discussed in this post, such as the law of the hammer, the over-reliance on data, the limits of extrapolation, the challenges imposed by interventional and counterfactual reasoning, the “hermeneutic” nature of these self-contained and isolated systems and their limited ability to transfer knowledge.
Proponents of deep learning continue to defend their position vigorously. But this perception that AI is hitting a wall is forcing a rethink on the prospects for inductive systems. Taking the pulse of the community and some notable unmet expectations, the AI researcher Filip Piękniewski argues that disillusionment, not strong AI, is the only inevitable outcome.
Much like Pearl, Marcus envisions a future that combines elements of explanatory knowledge with induction. “The right move today may be to integrate deep learning, which excels at perceptual classification, with symbolic systems, which excel at inference and abstraction.” This is undoubtedly true. Explanatory power is already deployed in the methodologies of machine learning, as in the selection of data, the choice of assumptions that bias learning algorithms, and the background knowledge used to “prime the pump” of induction.
But these are just observations from experts. You already know this. You know that reality frequently defies your assumptions of uniformity and foils your best laid plans. You don’t just observe a sunrise, you know why it rises. You’ve even unpacked the rich explanatory structure of quantum mechanics. You know that the most advanced inductive systems produce only pre-scientific knowledge, crude by the standards of our best scientific explanations.
And so you can answer for yourself: Will the gaps between weak and strong intelligence be crossed in predictable incremental steps or bold conjectural leaps?
In my estimation, all the impediments to strong AI are dwarfed by the instrumental fascination with inductive systems. Inductive systems like deep learning are powerful tools. It’s entirely understandable, even expected, that we should start with their practical uses. When heat engines were first invented at the start of the industrial revolution, the initial interest was in their practical applications. But gradually, over 100 years, this instrumental view gave way to a much deeper theoretical understanding of heat and thermodynamics. These explanations eventually found their way into almost every modern-day branch of science, including quantum theory. The tool gave us heat engines; explanations gave us the modern world. That difference is truly breathtaking.
AI will follow the same progression, from these first practical applications to a deep theoretical understanding of knowledge creation. I hope it doesn’t take 100 years. But when it happens, strong AI will follow. Inevitably.

Thanks for reading! If you liked it, please hit the applause 👏 so others find it. I occasionally blog about startups, healthcare and AI, if you’d like to Follow me and our startup journey. You can also connect with me on Twitter and LinkedIn to share ideas or leave a comment below👇.

Cartwright, N. (1983). How the laws of physics lie. Clarendon Press.
Deutsch, D. (1998). The fabric of reality. Penguin.
Hall, B. (2017). Induction. http://www.bretthall.org/blog/induction
Marcus, G. (2018). Deep learning: a critical appraisal https://arxiv.org/abs/1801.00631
Marcus, G. (2018). In defense of skepticism about deep learning. https://medium.com/@GaryMarcus/in-defense-of-skepticism-about-deep-learning-6e8bfd5ae0f1
Meehan, E.J. (1968). Explanation in social science; a system paradigm. Dorsey Press
Pearl, J. & Mackenzie, D. (2018). The book of why: the new science of cause and effect. Basic Books
Piękniewski, F. (2018). AI winter is well on its way. https://blog.piekniewski.info/2018/05/28/ai-winter-is-well-on-its-way/
Ribeiro, M.T., Singh, S. & Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier https://arxiv.org/abs/1602.04938
Salmon, W.C. (1967). The foundations of scientific inference. University of Pittsburgh Press.
Shmueli, G. (2010). To explain or to predict? Statistical Science, 25(3), 289–310. https://projecteuclid.org/euclid.ss/1294167961