GitHub Copilot AI pair programmer: Asset or Liability?

Automatic program synthesis is a long-lasting dream in software engineering. Recently, a promising Deep Learning (DL) based solution, called Copilot, has been proposed by OpenAI and Microsoft as an industrial product. Although some studies evaluate the correctness of Copilot solutions and report its...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	The Journal of systems and software Jg. 203; S. 111734
Hauptverfasser:	Moradi Dakhel, Arghavan, Majdinasab, Vahid, Nikanjam, Amin, Khomh, Foutse, Desmarais, Michel C., Jiang, Zhen Ming (Jack)
Format:	Journal Article
Sprache:	Englisch
Veröffentlicht:	Elsevier Inc 01.09.2023
Schlagworte:	Code completion GitHub copilot Language model Testing Language model GitHub copilot Code completion Testing
ISSN:	0164-1212, 1873-1228
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Automatic program synthesis is a long-lasting dream in software engineering. Recently, a promising Deep Learning (DL) based solution, called Copilot, has been proposed by OpenAI and Microsoft as an industrial product. Although some studies evaluate the correctness of Copilot solutions and report its issues, more empirical evaluations are necessary to understand how developers can benefit from it effectively. In this paper, we study the capabilities of Copilot in two different programming tasks: (i) generating (and reproducing) correct and efficient solutions for fundamental algorithmic problems, and (ii) comparing Copilot’s proposed solutions with those of human programmers on a set of programming tasks. For the former, we assess the performance and functionality of Copilot in solving selected fundamental problems in computer science, like sorting and implementing data structures. In the latter, a dataset of programming problems with human-provided solutions is used. The results show that Copilot is capable of providing solutions for almost all fundamental algorithmic problems, however, some solutions are buggy and non-reproducible. Moreover, Copilot has some difficulties in combining multiple methods to generate a solution. Comparing Copilot to humans, our results show that the correct ratio of humans’ solutions is greater than Copilot’s suggestions, while the buggy solutions generated by Copilot require less effort to be repaired. Based on our findings, if Copilot is used by expert developers in software projects, it can become an asset since its suggestions could be comparable to humans’ contributions in terms of quality. However, Copilot can become a liability if it is used by novice developers who may fail to filter its buggy or non-optimal solutions due to a lack of expertise. [Display omitted] •We investigate the quality of the code Copilot generates as an AI pair programmer.•Copilot provides efficient solutions; but some are buggy and/or non-reproducible.•Its solutions are more buggy but easier to fix compared to humans’.•Copilot’s suggestions are comparable to humans’ contributions in terms of quality.•Copilot can become an asset for experts, but a liability for novice developers.
ISSN:	0164-1212 1873-1228
DOI:	10.1016/j.jss.2023.111734