Source: Minsky, M., & Papert, S. (1973). Proposal to ARPA for Continued Research on A.I. for 1973 (MIT Artificial Intelligence Laboratory Artificial Intelligence Memo, Issue.
MASSACHUSETTS INSTITUTE OF TECHNOLOGY — A. I. LABORATORY
Artificial Intelligence Memo No. 284 | June 1973
Proposal to Arpa for Continued Research on A.I. for 1973
by Marvin Minsky and Seymour Papert
| GRANT & CONTRACT ACKNOWLEDGMENT Work reported herein was conducted at the Artificial Intelligence Laboratory, a Massachusetts Institute of Technology research program supported in part by the Advanced Research Projects Agency of the Department of Defense and monitored by the Office of Naval Research under Contract Number N00014-70-A-0362-0003. Reproduction of this document, in whole or in part, is permitted for any purpose of the United States Government. |
PREFACE
The Artificial Intelligence Laboratory proposes to continue its work on a group of closely interconnected projects, all bearing on questions about how to make computers able to use more sophisticated kinds of knowledge to solve difficult problems. This proposal explains what we expect to come of this work, and why it seems to us the most profitable direction for research at this time. The core of this proposal is about well-defined specific tasks such as extending the computer’s ability to understand information presented as visual scenes, or in natural, human language. Although these specific goals are important enough in themselves, we see their pursuit also as tightly bound to the development of a general theory of the computations needed to produce intelligent processes. Obviously, a certain amount of theory is needed to achieve progress in this and we maintain that the steps toward a comprehensive theory in this domain must include thorough analysis of very specific phenomena. Our confidence in this strategy is based both on past successes and on our current theory of knowledge structure.
Our proposed solutions are still evolving, but they all seem to revolve around new methods of programming and new ways to represent knowledge about programming. The field of Artificial Intelligence has made enormous progress in the past few years toward becoming a scientific subject. This proposal deals with our main goals, both for long-range research and for particular application areas. The general technical position of our 1970 and 1971 proposals still represents the direction of our approach. However, although those proposals are relatively explicit about the high-level problems motivating our research, they do not give a very clear picture of the actual projects, or of their practical consequences. In this proposal we shall concentrate on giving a more concrete view of our immediate goals. The next four sections are each directed to an “area” of application. It will be obvious that the concepts and methods used in these areas are closely related. Our intentions about the applications vary, naturally. In some cases we have major efforts toward completing effective operational prototypes. In the rest, we expect to produce demonstration prototypes, or only to develop some theory, clarify problems, and attempt to direct the attention of others to problems which we think need immediate attention.
PART I. SUB-PROJECTS AND MILESTONES
1.0 Anticipated Milestones
We define current and proposed sub-projects largely by relating them to milestone tests of achievement which we have set up to make the immediate goals and criteria of our work as unambiguous as possible. We do not try to predict the date of success for each milestone beyond saying that we would be surprised if any of them took as much as three years. It would be best to begin by discussing a milestone to whose achievement we are firmly committed because we are sure that it is in the state of the art and because it will have important theoretical and industrial consequences. It combines in a very interactive — more than merely additive! — way a large number of theoretically related sub-systems that have, for practical reasons, been built up separately.
1.1 The Language-Vision-Action Demonstration
This milestone will be passed when we can give the computer (equipped with eyes and hands) orders in brief, business-like language such as:
| SAMPLE COMMAND “I think there is a defective diode on this circuit board. Probably the third in the row on the left. Check and replace if necessary.” |
For this test to be meaningful, we must assume that:
- the computer has no prior knowledge about the particular circuit board (e.g. in the form of a diagram or symbolic description).
- the computer is able to see the circuit board by natural vision (so, for example, there are no special marks on it like the magnetic characters on checks).
- the English expressions used really are free (e.g. there is no hidden format beyond the restriction to what a real, not particularly smart, human technician would accept as a straight, natural and clear instruction).
The achievement of this demonstration depends on realizing a number of sub-goals of which some depend on parts of what is seen in the demonstration itself while others are more like the oil in the machine — no less vital and difficult for being invisible to the naive observer. These sub-goals include:
- extending the domain of the vision system from blocks with plane surfaces to the more natural objects found in electronic componentry with their curves, highlights, colors, stripes, shadows and so on.
- corresponding extension of the manipulative ability.
- extending the natural language programs in several directions: new domain (electronics instead of blocks); modalities such as probable-possible-necessary; modelling the person giving the instruction, so as to deal properly with such information as “I think…”
- further evolution of programming languages, computer control structures and debugging aids, in forms particularly suitable to work in Artificial Intelligence.
- theoretical problems in the area of representing knowledge, including “meanings” or “world models” constructed from visual and linguistic inputs.
This rough breakdown into sub-goals illustrates one of the major theoretical theses of our Laboratory. An “intelligence” is a complex system with a large number of interacting but separate parts. It can be understood and reproduced only by dealing specifically with these parts. We do not believe in the chimera of finding a single powerful “method” or “principle” which would give rise to intelligence the way the laws of Newton give rise to elliptic orbits. In surveying the variety of specific sub-projects mentioned below, the reader is entitled to be concerned about how many sub-problems we anticipate having to solve in order to achieve new levels of intelligence in automata! Is it tens? hundreds? millions? Are we merely scratching the surface? One of the purposes of unified milestone demonstrations (such as the one under discussion) is to show that the numbers can be contained. Speaking informally, we believe that the number of subsystems of intelligence is large, but not as large as the number of subsystems engaged in, say, an Apollo mission. The topics that have already been considered in the Laboratory constitute a significant fraction of what is necessary.
1.2 Other Milestones: Vision and the Blocks World
We will describe the other milestones more briefly, though they are no less serious. Some of them are easier and will be achieved earlier than the Language-Vision-Action demonstration in electronics; some are much more difficult.
The BLOCKS WORLD has played a central role as a culture medium for ideas about vision and will continue to do so despite the new concentration on other more useful and complex areas of work. To place new milestones in perspective we recall briefly some past history:
- The first important steps toward the present approach were the work, more than a decade ago, by H. Ernst on touch-sense controlled manipulation and by L. Roberts on machine scene-analysis of photographs of three-dimensional plane-surfaced objects.
- The next significant step, this one under ARPA support in the A.I. Lab (then part of Project MAC) was the development of visually-controlled tower-building programs in 1965–1966. The major sub-system was direct “natural” vision of real, but unoccluded and clearly illuminated blocks.
- In more recent milestone demonstrations, we have seen machine manipulation of several blocks in visually occluding relationships. Behind this practical step was a more important theoretical one, first taken by Guzman, of basing scene analysis on general knowledge about “bodies” or “objects” without using (as was done before) specific knowledge about particular kinds of bodies.
In this older work on vision, shadows and other side effects of lighting were treated as embarrassing complications to be minimized. However, these effects are sources of information, and we now understand them well enough to exploit them as such. A milestone for the very near future will be a Vision System to see highly shadowed and badly illuminated scenes, with the help of (rather than in spite of) the shadows. In the Language-Vision-Action demonstration, the relation between the hand and its shadow should be exploited to anticipate the contact of the hand and its target.
The MIT VISION SYSTEM is, at present, still an experimental prototype developed to explore new ideas both about machine vision and new styles of programming. It is the most powerful system available within its narrow domain of analysing monocular, polyhedral scenes. We are extending its capabilities to deal with more natural objects. Up to now, this system could deal only with polyhedral objects having uniform plane surfaces, such as blocks, wedges, pyramids, and the like.
The Vision programs generally use information of just one kind: light intensities on a two dimensional projection. An important imminent advance is the use of multiple sources of information. Some of these will be other dimensions of vision — color, range, etc. Other sources will be symbolic, in particular those derived from interaction with SHRDLU. We do NOT see as a significant milestone merely using top level commands in English, passed on through SHRDLU, to operate a Vision System. This may be a good demonstration to the outside world of the flexibility and utility of SHRDLU, but it is really quite easy. A much more significant step will be to marry SHRDLU and VISION so as to permit statements made in English to give useful advice to the vision program: “Look more carefully in the shadow of the cube to the right of the tall block”. A much more difficult step — sufficiently significant to count as a separate milestone if done in an insightful way — is telling the program in English how to extend its mini-world of competence. For example, one might tell a program designed to see and manipulate blocks about a good strategy for putting things in boxes.
1.3 Vision Outside the Blocks World
- Liquids: An excellent test-bed for ideas about changing shapes. The milestone is an eye-hand system that will make a cup of regular instant coffee.
- People-Watching: Our own research on the control and acquisition of motor (and postural) skills in humans will benefit from a program capable of observing a person engaged in a task such as learning to walk a tight-rope. A simple application of this is detecting when a person approaches dangerously close to the edge of a platform, and warning him.
- Interpretation of Drawings: The interpretation, by machine, of not-quite realistic or conventional drawings poses problems intermediate between vision and language. It is a task that has not yet been performed significantly by machine, yet seems ripe as a future milestone. We hope, through work like that of Goldstein, to develop a program that can describe in words — that is, by understanding — what is shown in a drawing, cartoon, or action-sketch with stick-figures.
1.4 Understanding English
Our Laboratory has pursued over the years a working hypothesis that can be stated briefly as: before one can get a machine to understand English, one must find how to make it understand at all. Translated into concrete terms, this means making programs to understand complex English statements about a simple mini-world which the computer is able to understand very thoroughly by drawing on a stack of special knowledge. An early milestone whose achievement encouraged this point of view was the early program, STUDENT, by Dan Bobrow. The latest clear break-through is SHRDLU by Terry Winograd.
A secondary milestone is the operational use of SHRDLU as a “front-end” to communicate with programs written for quite different, practical purposes. An early example of this kind of use is the coupling of SHRDLU to a question answering program made by C.C.A. We know of several other such projects. Though it is encouraging to see Artificial Intelligence programs being used more and more extensively, these uses of SHRDLU do not go beyond its original formal capacity. More substantial milestones which we see as significant and yet not far off include:
- Using More Linguistic Information: A program that makes significant use of tenses, modalities (e.g. possible-probable) and indirect reference (e.g. He thinks that…) would be a considerable advance.
- Getting Away With Less Linguistic Information: We have in mind using “common sense” interpretation to fill in only partially specified information.
- Extending SHRDLU in English: A most important milestone will be passed when we are able to describe, in English, extensions to SHRDLU itself. SHRDLU’s interpretational power would then act in a bootstrapping fashion.
1.5 Computers, Knowledge, and Intelligence
The world of Science is well-organized for the management of some kinds of knowledge. Its theories provide a firm (though not, of course, infallible) understanding of what kinds of knowledge are relevant to such tasks as predicting astronomical events or designing bridges. There are institutionalized repositories of such knowledge (handbooks, tables, encyclopedias) as well as the means of transmitting it to the next generation (schools).
Not all knowledge has been treated in this formal manner. In particular, a very large body of knowledge has traditionally been neglected for the simple reason that “everyone knows it anyway”. This is the kind of common sense knowledge that leads one, for example, to rearrange the contents of a box (or throw something out) when more space is needed.
Research in Artificial Intelligence is forced to deal with this kind of knowledge quite explicitly. One might argue that the “intelligent computer” ought to acquire such knowledge by the tacit, informal process that leads humans to have it without explicit formalization. This is certainly true in some sense; but even to understand what is being asserted we need to formalize more explicitly, and characterize more insightfully, the kinds of knowledge in question and the kinds of processes that might lead to its acquisition.
Among our (human) ways to acquire knowledge, two stand out beyond others and our work has centered on them: Language and Vision. The rest of this proposal is divided into three parts, accordingly:
- Part II is concerned with issues related to logic, general knowledge, and common sense, in a context centered mainly around problems of understanding natural language. This section also includes closely related work on understanding procedures, programming languages, and descriptions of programs.
- Part III discusses issues connected with robotics, machine vision, manipulation, and other activities that involve physical-world interactions.
- Part IV discusses work needed to develop tools: hardware, programming languages, software, and other things needed to support the goal-directed projects.
PART II. LOGIC, GENERAL KNOWLEDGE, AND COMMON SENSE
A number of projects in the Laboratory are centered around the problems of understanding natural language. These studies include both theoretical and practical problems:
- Section 2.1 — Polishing the Language-Understanding Program: Terry Winograd and group
- Section 2.2 — More Powerful Problem-Solving in SHRDLU (CONNIVER): Winograd, W. Martin, G. Brown
- Section 2.3 — New Models for Meaning: Minsky, Miller, Charniak, and others
- Section 2.4 — New Grammar Organizations: D. McDonald, A. Rubin, V. Pratt
- Section 2.5 — Theories of Generic Expressions and Typical Situations: R. Moore, M. Marcus
- Section 2.6 — Learning to Understand (TOPLE): D. McDermott
- Section 2.7 — Understanding Stories: E. Charniak
- Section 2.8 — Sussman’s Program: HACKER: G. Sussman
- Section 2.9 — Knowledge About Procedures and Their Likely Bugs: G. Sussman
- Section 2.10 — Goldstein’s ‘Solution’ of the Halting Problem: I. Goldstein
- Section 2.11 — Goldstein’s Program-Understanding Program: I. Goldstein
2.1 Polishing the Language-Understanding Program
A project that will be completed this year is the conversion of Winograd’s language-understanding program (called SHRDLU) from a first demonstration prototype into a generally accessible, well-documented experimental facility for further development of linguistic models and applications that are related to the basic theory behind this model. The new system contains extensive tracing and debugging aids, and a detailed manual is being prepared. This will make the system accessible to users who begin with very little prior knowledge about its details. The system is available over the ARPA network.
2.2 More Powerful Problem-Solving in SHRDLU
Several students are rewriting parts of the language understanding system to use CONNIVER instead of Micro-Planner. The new language, which permits cross-reference between different parts of a problem-solving process, will provide a better base for extending the language program to a larger world that can use hypothetical contexts, plausible reasoning, and more complete handling of tenses.
We are also working with other groups in their uses of this language understanding system. The project of W. Martin, at Project MAC, will probably use it as a means of communication with their automatic programming systems, and G. Brown, also at MAC, is exploring the possibilities of basing a language translator on SHRDLU; at present she is working on a very small specific set of problems involving German noun cases and prepositions, but the basic ideas may be generalizable.
2.3 New Models for Meaning
At a theoretical level, several of the staff and students are exploring some new formalisms for representing meanings. One problem is to tie together a variety of phenomena that have been addressed more or less separately by such models as Fillmore’s Case Grammars, Abelson’s ‘molecules, texts, scripts, etc.’, Schank’s Conceptual Grammars, Halliday’s Systemic Grammars, etc. We believe that these can be tied together in a more consistent way by viewing language understanding as an active procedure, constantly engaged in an operation of ‘fitting together’ inputs and the implications of inputs. These theories are just beginning, but we hope they will lead to a second generation system much better equipped to handle plausible reasoning, incomplete inputs and a wider range of ways in which language conveys meaning. Among their new elements are what might be called ‘scenarios’ — structural elements that describe ‘the ways things usually happen’ — that are substantially larger than the kinds of elements found in earlier syntactic or semantic theories.
2.4 New Grammar Organizations
David McDonald is working on problems of generating sentences that represent meanings in ways that are responsive to the needs of the system’s user. This means that the system must use its knowledge about what (it thinks) the user believes. McDonald’s approach to generating coherent discourse views it as concerning multiple procedures operating on a common data structure. First, there is a logical process concerned with choosing words and producing structures which convey the underlying meaning. Second, there is a discourse procedure whose job is to structure the output for coherence, making the necessary inter-sentence connections and introducing pronouns when appropriate. Third, there is a syntactic specialist to make sure the output is in a proper grammatical form. McDonald wants a heterarchical system in which these time-share and interact cooperatively.
A. Rubin is working on a new version of Winograd’s grammar, exploring different types of organization which get away from the highly linear organization of the current grammar. Words such as “a” and “the”, or prepositions and question words have immediate syntactic implications and should be able to direct the parsing process. Rubin is also studying problems of extending the system to use complex tense and time semantics — perfect and progressive tenses, futures, and modals.
2.5 Theories of Typical Situations and Generic Expressions
R. Moore has been exploring some of the problems in connecting English expressions to their underlying logical forms. He is concentrating particularly on English quantifiers and their connections with mathematical logic. A word like “any” can have several different meanings in logical terms. In “Did anyone come to see me?” it corresponds to the existential quantifier (∃X), but in “Intercept anyone who comes to see me” it represents a universal quantifier (∀X).
M. Marcus is studying “generic” nouns and their role in commonsense reasoning. When one says “A bird can fly”, the traditional logical representation (∀X)(Bird(X) → Fly(X)) is inadequate because ostriches and birds with broken wings cannot fly. One usually means “If X is a typical bird, X can fly”. Understanding what is “typical” depends on detailed general knowledge about the domain rather than on discovering a new formal logical quantifier.
2.6 Learning to Understand (TOPLE)
D. McDermott is developing a program called TOPLE (written in CONNIVER) which tries to understand simple declarative statements about a blocks-like world. TOPLE can “visualize” spatial relations, guess causal explanations, and make predictions about the future course of events. McDermott’s work addresses the problem of translating natural declarative statements into belief systems that record reasons for beliefs and handle abbreviated or ambiguous inputs through plausible reasoning.
2.7 Understanding Stories
E. Charniak continues to work on text-comprehension through exhaustive analysis of children’s story fragments. He investigates how knowledge is accessed via “demons” or semi-autonomous programs activated by key words or events. For example, resolving pronoun references in stories about birthday presents requires active rules about social expectations (e.g., returning unwanted gifts to a store).
2.8 – 2.9 Sussman’s Program: HACKER and Procedural Bugs
Gerald Sussman’s program HACKER explores learning from experience through automatic debugging. Rather than storing purely domain-specific rules, HACKER represents knowledge about general programming strategies and bug patterns. When a program fails during execution (e.g. attempting to move a block that has another block sitting on top of it), HACKER diagnoses the underlying procedural interaction bug and automatically patches the program code using generalized operations such as ‘protection’ (inhibiting changes to necessary preconditions of super-goals).
2.10 – 2.11 Goldstein’s Program-Understanding Program
Ira Goldstein is designing a monitor (SUPERMONITOR) capable of analyzing programs written by novice programmers or automated program generators. Although the general halting problem is formally undecidable, real-world infinite loops arise from recognizable structural flaws (e.g., recursive calls without stop rules or failing parity conservations when counting down by twos). Goldstein’s system also analyzes graphic programs written in LOGO (such as drawing stick-men) to diagnose semantic bugs and understand procedural intentions.
PART III. COMPUTER-CONTROLLED VISION AND MANIPULATION
Research on computer vision and its applications is entering a new phase in the Laboratory. In the past, the work centered around prototype problems in the BLOCKS WORLD. Current research is pointed at the problems of natural environments within a modular “heterarchical vision system”.
- Section 3.1 — Heterarchical Vision System and Project Direction: P. Winston, B.K.P. Horn
- Section 3.2 — Generalization of Labelling Theories: D. Waltz
- Section 3.3 — Grouping and Tactile Scene-Analysis: T. Finin, T. Lozano-Perez
- Section 3.4 — Analysis of Curved-Line Drawings: M. Adler
- Section 3.5 — Color Vision: M. Lavin
- Section 3.6 — Touch and Tactile Programming: D. Silver
- Section 3.7 — ‘Low-Level’ Vision Programs: J. Lerman
- Section 3.8 — Physical Knowledge and Stability: S. Fahlman
- Section 3.9 — Liquids and Hand-Eye Tracking: R. Woodham
- Section 3.10 — Electronic Assembly: Another Robot ‘World’: P. Winston and others
- Section 3.11 — Analyzing Complicated Objects: J. Hollerbach
- Section 3.12 — Groups, Descriptions, and Conflicts: M. Dunleavy
3.1 – 3.2 Heterarchical Vision and Labelling Theory
Prof. P. Winston supervises scene-analysis theory, automatic learning of 3D structures, and system heterarchy, while B.K.P. Horn supervises image processing systems. David Waltz’s recent thesis expanded line-labelling semantics in polyhedral vision, demonstrating that geometric junction constraints (concave, convex, obscuring edges, cracks, and shadow boundaries) severely restrict allowable interpretations, enabling rapid constraint propagation.
3.3 – 3.5 Grouping, Curves, and Color
T. Finin and T. Lozano-Perez are extending grouping heuristics to highly occluded scenes, integrating tactile probes to confirm or reject structural hypotheses without causing physical collapse. M. Adler is extending line-labelling to curved-line drawings and cartoon conventions. M. Lavin is investigating physiological vs. spectral color vision theories to improve robot perception in non-uniform lighting environments.
3.6 – 3.9 Tactile Sensing, Stability, and Liquids
D. Silver utilizes a six-axis force-sensing wrist for tactile tasks like turning cranks and threading nuts. S. Fahlman is completing a construction-planning specialist that reasons about gravity, balance, scaffolding, and counterweights. R. Woodham is developing a liquid-pouring specialist in a coffee-making environment, requiring real-time visual feedback to track rising liquid levels and dynamically adjust hand-eye manipulation.
3.10 – 3.12 Electronic Assembly and Complex Objects
The laboratory is expanding robot environments to electronic assembly and inspection (orienting components, soldering, identifying color-coded resistor stripes). J. Hollerbach is developing multi-level structural descriptions for complex polyhedra, bottles, and cups. M. Dunleavy is investigating architectural layout heuristics (walls, doors, chimneys) to resolve local geometric conflicts in repeated structural groups.
PART IV. DEVELOPMENT OF RESEARCH METHODS AND TOOLS
4.0 – 4.1 Programming Languages for A.I. and PLANNER
A major impediment to progress in A.I. has been the lack of appropriate programming languages. Building such languages is equivalent to developing a formalism and fundamental primitive concepts for A.I. Carl Hewitt and collaborators are developing a modular activation formalism that unifies PLANNER-like control structures, enabling efficient implementation, formal control structure analysis, and feasibility studies for specialized hardware processors running 40–60 times faster than a PDP-10.
4.2 – 4.5 Parsing, Debugging, and Theorem Proving
Prof. V. Pratt is developing LINGOL, a linguistics-oriented language implementing efficient context-free parsing with user-supplied semantic evaluation code. S. Markowitz is developing advanced debugging tools for pattern-directed invocation languages embedded in LISP. E. Freuder is reorganizing vision systems around heterarchical control. In automatic theorem proving, A. Brown, A. Nevins, and J. Geiser are investigating heuristic derivation in abstract group theory, case analysis, and reasoning in empirical data bases containing potential inconsistencies.
4.6 Education and Cognitive Development
Seymour Papert and the LOGO group explore the symbiosis between Machine Intelligence and Natural Intelligence in children. The procedural approach to thinking translates directly into educational practice — replacing abstract formal logic with intuitive, qualitative physical reasoning and debuggable computational structures.