AI

AI: What We Should Understand and How We Might Develop

Facing the present and looking toward the future

Posted by Bruce Lee on 2023-07-25

About Me

Welcome to my blog! This is where I collect my observations and notes on programming and technology. The main subjects range from implementation details to broader ideas about programming.

Main Topics

  • Engineering Projects: Exploring implementation details and how technical systems work.
  • C/C++: Notes on language features and programming techniques.
  • The Programmer’s Perspective: Ideas about developing a career and a way of thinking as a programmer.

For more, visit the categories page.

Contact

If you have questions or would like to discuss something, please get in touch through the About page.

Thank you for reading and for your support. I hope these notes help you on your own technical journey!


This essay records the author’s perspective in July 2023. Its employment forecasts, salary expectations, and proposed career categories are opinions from that period. Technical clarifications are marked separately so that those predictions are not mistaken for established facts.

A changing industry

The IT industry I knew was associated with rapid expansion, large internet-company revenues, and enormous demand for programmers. A substantial business might require a team covering databases, interfaces, third-party APIs, payments, and small applications.

My concern in 2023 was that AI could allow smaller teams to deliver some of that work, changing ordinary programming jobs and putting pressure on wages or staffing. This is a forecast about selected tasks and businesses, not evidence that one AI scientist can replace every software team.

I also saw a contrast in product economics. Many earlier internet products attracted users with free access or subsidies, then monetized a large audience through advertising and paid services. Some AI products charged from the beginning, perhaps after a few trial uses. These are business patterns, not rules that apply to all products.

What is behind ChatGPT?

The 2017 paper Attention Is All You Need introduced the Transformer architecture and became an important foundation for later language models. The familiar diagram is a useful starting point:

Transformer encoder and decoder

The encoder takes an input sequence A = {a₁, a₂, …, aₙ} and produces contextual representations B = {b₁, b₂, …, bₙ}. Tokens are represented numerically, and attention relates positions to one another. Position information matters: changing the order of the words in a sentence can change its meaning.

The decoder in that encoder-decoder design uses the encoder output together with the preceding output tokens. Its shifted-right input allows it to predict the next element from what came before. Autoregressive generation repeatedly computes a conditional distribution for the next token and selects or samples a token, extending the sequence.

The original explanation described choosing the most likely token every time. That is one decoding strategy; sampling can also be used. Tokenization is not necessarily character-by-character: tokens can be whole words, word fragments, characters, or byte-related units depending on the tokenizer.

BERT and GPT are both associated with Transformer architectures, but use them differently. BERT is encoder-oriented, while GPT-style language generation is decoder-oriented. The original claim that every AI model is an autoregressive Transformer is incorrect; AI includes many other architectures and learning methods.

What does N× mean?

The repeated blocks in the diagram indicate a stack of layers. Depth can increase a model’s capacity, while also increasing computation and resource requirements. The right size depends on data, task, training, and engineering constraints.

The numbers six layers and a representation width of 512 in the original discussion come from the original Transformer’s base configuration. They should not be read as specifications for every OpenAI model or for ChatGPT. A representation width is a vector dimension, not simply the number of decimal digits needed to encode a token.

What might an algorithm engineer do?

A company adopting AI needs people who understand research, evaluate methods, run experiments, prepare data, and adapt systems to its tasks. Reading papers and trying model configurations are parts of that work. Layer counts and representation widths are examples of architectural choices, but routine fine-tuning of an existing model usually does not mean freely changing those dimensions.

My emphasis in the original essay was on adaptation to specialized domains. A general model may not meet the needs of legal, medical, counseling, or code-obfuscation work without additional data, evaluation, and system design. Here vertical domain means a narrow professional or industry context. The same question can require different terminology, evidence, and constraints in different professions.

Fine-tuning is one way to adapt a model. It is not the entire job, and it does not by itself guarantee reliable professional answers. Retrieval, prompting, software integration, evaluation, and human review can also matter.

Why build on existing models?

The essay’s argument was that strong pretrained models create opportunities for many engineers to adapt them, while a smaller group works directly on foundational research. Training a new large model from scratch requires resources that most application teams do not have.

Looking back from mid-2023, I associated the release of conversational interfaces and the progression from GPT-3 through later systems with a striking public shift. The source contains conflicting statements about which transition was most impressive. Its broader point is the perceived importance of both stronger models and making them useful through a conversational product.

Technical clarification: the development of those systems cannot be reduced to a claim that all progress was merely fine-tuning an unchanged base. The original essay did not establish their complete training methods or architecture differences.

How should we choose a direction?

I worried that routine programming would be among the first kinds of work affected. To reason about that, I proposed my own division into deterministic and non-deterministic fields, borrowing language I had encountered in a robotics course.

By deterministic, I meant tasks with well-defined conditions and predictable evaluation. By non-deterministic, I meant work whose outcome or interpretation remains open to uncertainty. This was an informal career heuristic, not a mathematically rigorous classification of whole industries.

Go was an early example because its rules and game state are precisely defined. The source’s explanation that Zero in AlphaGo Zero means unbeatable is incorrect; the name refers to learning without human game examples in the method described by its authors.

I listed manufacturing, parts of driving, financial risk analysis, fraud detection, programming, reverse engineering, statistics, and control systems as areas where AI could handle increasingly formalized tasks. These examples need qualification. Real-world driving and finance contain uncertainty, distribution shifts, and imperfect information; they are not deterministic merely because an algorithm is involved. Nor does the essay establish that AI always makes better judgments than people in them.

Translation made the limits of my division visible. Nuance, emotion, and context can require human interpretation, yet automated translation had already achieved useful results through extensive research, data, and human evaluation. The original suggestion that AI as a whole began with translation is too narrow.

My practical conclusion was to learn to collaborate with AI and to value work involving judgment and ambiguity. Even in fields that are not fully predictable, AI can assist. Conversely, strong expertise does not make a person categorically immune to technological or economic change.

What kind of programmer might remain valuable?

My answer was a combination of broad software engineering ability and AI knowledge. I argued for solid foundations in operating systems, assembly, compilers, networks, Linux, calculus, discrete mathematics, and linear algebra, together with the ability to connect front-end and back-end work.

As a computer science and engineering student, I saw those foundations as the result of sustained study. I expected it to become harder to enter the industry through a short course covering only databases and basic coding. That expectation reflects my perspective, not a rule that people changing careers cannot build the same competence.

The original essay uses strong language about most programmers being replaced. The underlying concern is a shift in the mix of valuable tasks, but it supplies no evidence for a specific replacement rate.

The next wave of adaptation work

In 2023 I expected more companies to want models adapted to their own data and workflows. I called these private models, grouping together several different ideas: customized behavior, use of proprietary data, and privately deployed systems.

Those distinctions matter. Accessing a hosted API does not automatically produce a private deployment or a separately owned model, and API availability is not the same thing as an open-source license. Building on an already adapted model can reduce work, but suitability depends on the task and deployment requirements.

I expected demand for engineers who could prepare data and fine-tune or integrate models, and thought existing programmers with strong foundations would be well positioned to learn. For students, my suggested path was mathematics and computing fundamentals followed by deeper AI study; for people already employed, workplace training could provide an entry point.

I also speculated that competition and credential requirements might increase as more people acquired similar technical skills. The comparison with finance was about barriers to entry, not a claim that the industries have identical hiring practices.

Salary expectations and learning

The original essay was optimistic about a wave of well-paid AI engineering jobs, even suggesting annual compensation above one million yuan for some successful transitions. It compared that hope with examples of earlier IT salaries in the hundreds of thousands. These are speculative expectations and anecdotes, not a promise or a reliable salary forecast.

Likewise, describing fine-tuning as merely guessing probabilities understates the difficulty of data quality, optimization, evaluation, deployment, and reliable system behavior. The useful personal goal remains to become a programmer who can use AI thoughtfully and understand its limitations.

A more personal model

One possibility that interested me was a model reflecting an individual’s writing style and habits while retaining general knowledge. Such a system might help someone view their behavior from another perspective or give family and friends a familiar conversational style. I also imagined people speaking with a model after the person it represented had died.

That is a proposed direction rather than a claim that a model reproduces a person’s consciousness or reliably represents their wishes. Consent, privacy, and clarity about the simulation would matter.

The future I imagined was full of change. My hope was that readers would find a field they enjoy, build the knowledge to contribute to it, and learn how to work with the tools that emerge.


If you like this blog or find it useful for you, you are welcome to comment on it. You are also welcome to share this blog, so that more people can participate in it. All the images used in the blog are my original works or AI works, if you want to take it,don't hesitate. Thank you !