A Deep Dive into rasa
With the rise of personal assistants like Siri, Google Assistant, Cortana, Alexa and how they’re all called, there seems to be a great interest in chat bots, which are basically small text-based Alexas! So - exactly like 20 years ago - many companies want cool chat-bot interfaces for Facebook, WhatsApp or simply their website. That’s where rasa comes into play. rasa is a startup that provides a framework for building bots utilizing the hottest and newest approaches, including Machine Learning™.
What is rasa?
That is a difficult question. It is the startup, the “platform”, the “stack”, the framework. So basically it is piece of software that you can use to build chat-bots and also a company that uses that framework and other tools (as far as I understood it) to build chat-bots for companies.
The software itself, which is what I will focus on in this blog entry, is split into two projects: rasa core and rasa nlu.
rasa nlu handles the natural language understanding. It takes the sentences typed by the users, classifies it into one of several intents (i.e. what is the user intending to do?) and detects entities that are mentioned. So for example, when a user asks: “What is the weather going to be tomorrow in Berlin?” a well trained NLU could return the intent weather_request with the entities date: "tomorrow" and location: "Berlin". rasa nlu uses spacy under the hood and does its job very well (in my experience).
The second part, rasa core, is what I will focus on in this blog post. It is basically “the rest” (which is probably why it’s called core) and it is mainly to solve the task “What does the system should say/do given the dialogue up until now”. For this, it uses a Recurrent Neural Network that gets the history of the last actions that were taken by the user as well as the system and predicts what action should be taken based on that.
Just to show you how hip this framework is, let me tell you that the training data (dialogues that are used for the RNN to train on) are stored in a Markdown file and the domain is specified in YAML. But in all seriousness, this helps to remove the barrier for newcomers that want to build their own bot.
I recommend you to at least look at their basic tutorial, where you can see that you don’t have to write a single line of code to build a (very basic) bot. I think that is very impressive (of course if you want a bot that actually does something, it gets slightly more complex).
How does it work?
Damn. This is also a difficult question. In principle, the tutorials on the rasa website try to let you create bots by only showing you the parts that you really need to know. Which is good if you just want to build a bot quickly, but is not good for the deeper understanding.
Throughout the tutorials one page is always linked when it gets to the interesting “under the hood” stuff. The page is titled “Plumbing - How it all fits together” and mostly consists of this image:
Sadly, this is not that informative, but let me still try to explain this image a little bit. Basically, chat-bots are a pipeline: message comes in, bot does stuff, message comes out. In this case, the message arrives at the Interpreter 1. The Interpreter is rasa nlu and as I described above, it converts the text into something meaningful for the computer, namely an intent and entities.
This information is then handed to the Tracker 2. The Tracker is basically the control unit of the chat bot. It keeps track of what the system and user has already said, what information was given by the user (slots that are filled). The Tracker takes the new dialogue act together with the acts from the last few turns and hands them to the Policy 3. The Policy is the RNN that then determines what Action 4 the bot should take. The Action does its thing (for example retrieving the weather from an API), updates the Tracker (so that the Tracker knows which Action was being executed and to update the state accordingly) and sends a message to the user 6.
That seems very simple and bots with only a few stories (i.e. training data) already work rather well. So, that’s it, right?
What’s a Memoization?
Digging deeper into the rasa core source code, I saw that this image is oversimplifying (duh). In the second tutorial I found the following code:
agent = Agent(domain_file, policies=[MemoizationPolicy(), RestaurantPolicy()])
This code is executed during training and it seems the the agent is given a list of policies. The RestaurantPolicy is a policy created in the tutorial and is basically an LSTM RNN. The MemoizationPolicy however is not really explained. Here is what the documentation says:
The link to the Plumbing page unsurprisingly did not yield new insights. However, the picture with the one policy is not quite complete.
Looking at the source code for the tracker I could find the PolicyEnsemble, a class that can incorporate different policies (i.e. decision makers). So, at each step of the dialogue multiple policies are executed and the best one gets the bid.
class SimplePolicyEnsemble(PolicyEnsemble):
def __init__(self, policies, known_slot_events=None):
super(SimplePolicyEnsemble, self).__init__(policies, known_slot_events)
def probabilities_using_best_policy(self, tracker, domain):
result = None
decision_maker = None
max_confidence = -1
for p in self.policies:
probabilities = p.predict_action_probabilities(tracker, domain)
confidence = np.max(probabilities)
if confidence > max_confidence:
max_confidence = confidence
result = probabilities
decision_maker = p
logger.debug("%s made the decision!" % decision_maker)
return result
The SimplePolicyEnsemble asks each policy for their probabilities given the tracker (i.e. the current state) and the domain and selects the decision of the policy with the highest confidence. Note that I added a debug output (decision_maker) which prints the policy that had the highest confidence, just so I could see in each turn which policy was responsible for the actions of the bot.
Let’s see what action probabilities the MemoizationPolicy is producing:
def predict_action_probabilities(self, tracker, domain):
x = self.featurize(tracker, domain)
tracker_state = ["{}".format(e)
for e in self.featurizer.decode(x,
domain.input_features)]
logger.debug('Current tracker state [\n\t{}]'.format(
"\n\t".join(tracker_state)))
memorised = self.recall(x, domain)
result = [0.0] * domain.num_actions
if memorised is not None and self.is_enabled:
logger.debug("Used memorised next action '{}'".format(memorised))
result[memorised] = 1.0
return result
The MemoizationPolicy recalls, if it saw the current dialogue in the training data. If so, it returns what’s in the training data with 100% confidence. That means if you stay in the “golden path” (the stories that were trained) the memoization policy does exactly what’s in the training data.
Who is pulling the strings?
With the added debug output in the policy ensemble and the memoization policy, I saw that most of the time the memoization policy decided the next turn. In it self, this is not a bad thing. If a dialogue is exactly like in the training data, it is probably not a bad idea to do what the training data says.
However, when I disabled the memoization policy the neural network did not perform as expected.
Input:
> hi
Logging:
2018-04-26 08:54:11 DEBUG rasa_core.processor - Received user message 'hi' with intent '{'name': 'greet', 'confidence': 0.9174638293066442}' and entities '[]'
2018-04-26 08:54:11 DEBUG rasa_core.policies.ensemble - <__main__.RestaurantPolicy object at 0x12d3c6438> made the decision!
2018-04-26 08:54:11 WARNING rasa_core.processor - Circuit breaker tripped. Stopped predicting more actions for sender 'default'
Output:
--> I'm on it
--> I'm on it
--> I'm on it
--> what kind of cuisine would you like?
--> I'm on it
--> I'm on it
--> I'm on it
--> what kind of cuisine would you like?
--> I'm on it
--> I'm on it
That’s not looking good. To clarify: I removed some logging messages, but in principle the bot is producing more and more actions, more and more responses until the rasa core activates a “circuit breaker” so that the bot is not stuck in an endless loop. The rasa nlu classified the input correctly (intent greet and no entities mentioned) and the Restaurant policy made the decision.
Why is this happening? Short answer: action_listen. Long answer: there is a “hidden” action the bot can do: the listen action. After the user writes something the policy is activated again and again until it produces an action_listen where the bot awaits new input from the user. This way the bot is able to answer with more than one action. For example, if the user has requested to search for a restaurant, the bot can execute the action action_on_it - telling the user that it could take a while - and then it could execute the bot.ActionSearchRestaurants action.
That’s a neat feature, but it seems to make problems with the neural network. The network (called the KerasPolicy) is asked what to do next. It predicts to say “I’m on it”, sends the message to the user and informs the tracker. The tracker then takes the update dialogue and activates the policy again. Because there is something wrong with our neural net, it again predicts the action_on_it action and round and round it goes.
One other noteworthy case that seems to appear often is that after a users input the machine learning policy immediately predicts the action_listen. That way the bot simply is silent and awaits new inputs from the user.
What’s next?
So the machine learning approach is not that robust and useful as it seemed! I don’t want to say that it is not working, just that the parameters included in the tutorial are clearly not working out.
I will try to tweak some parameters and see if I can get the system to run in a more acceptable way. I will also look into the featurization of the dialogue state. Maybe there are some insight to why the current settings are not working.


