[D] How did Microsoft’s Tay work?
How did AI like Microsoft’s Tay work? This was 2016, before LLMs. No powerful GPUs with HBM and Google’s first TPU is cutting edge. Transformers didn’t exist. It seems much better than other contemporary chatbots like SimSimi. It adapts to user engagement and user generated text very quickly, adjusting the text it generates which is grammatically coherent and apparently context appropriate and contains information unlike SimSimi. There is zero information on its inner workings. Could it just have been RL on an RNN trained on text and answer pairs? Maybe Markov chains too? How can an AI model like this learn continuously? Could it have used Long short-term memory? I am guessing it used word2vec to capture “meaning”
submitted by /u/RhubarbSimilar1683
[link] [comments]