1 00:00:05,239 --> 00:00:07,259 [music] 2 00:00:09,825 --> 00:00:11,845 [music] 3 00:00:15,440 --> 00:00:17,119 And 4 00:00:17,119 --> 00:00:20,000 for our next talk, I'd like to welcome 5 00:00:20,000 --> 00:00:23,600 to the stage Marty Kilishek. 6 00:00:23,600 --> 00:00:26,320 Please make him most welcome. [applause] 7 00:00:26,320 --> 00:00:29,560 Thank you. 8 00:00:30,130 --> 00:00:30,720 [sighs] 9 00:00:30,720 --> 00:00:34,960 Hello. Hello. Uh, my name is Marty and I 10 00:00:34,960 --> 00:00:38,160 lead the advanced engineering team of 11 00:00:38,160 --> 00:00:41,200 Incubator at NBNo. And for the past 10 12 00:00:41,200 --> 00:00:43,280 to 11 months, we've been experimenting 13 00:00:43,280 --> 00:00:47,039 with agents. Now, 14 00:00:47,039 --> 00:00:49,280 we've had some fun with agents and we've 15 00:00:49,280 --> 00:00:51,039 also had some challenges with agents 16 00:00:51,039 --> 00:00:53,840 that I'm sure you all would appreciate. 17 00:00:53,840 --> 00:00:55,760 And what would have helped us out at the 18 00:00:55,760 --> 00:00:57,199 time when we started which was roughly 19 00:00:57,199 --> 00:01:00,079 October last year was if we had agent 20 00:01:00,079 --> 00:01:03,760 skills. So this talk is about before 21 00:01:03,760 --> 00:01:07,040 agent skills emerged and then after. So 22 00:01:07,040 --> 00:01:09,200 I trust that at the end of this talk for 23 00:01:09,200 --> 00:01:10,720 those individuals that are currently 24 00:01:10,720 --> 00:01:13,360 developing agents or using agents to 25 00:01:13,360 --> 00:01:15,280 have a methodology for how you proceed 26 00:01:15,280 --> 00:01:16,960 forward through this non-deterministic 27 00:01:16,960 --> 00:01:20,799 world. Uh, first off, NBN. For those 28 00:01:20,799 --> 00:01:22,720 that are not familiar or who would like 29 00:01:22,720 --> 00:01:27,200 a reminder, NBN is Australia's broadband 30 00:01:27,200 --> 00:01:30,320 access network and it is the access 31 00:01:30,320 --> 00:01:32,240 network that your internet service 32 00:01:32,240 --> 00:01:35,439 providers sell you access to. It the 33 00:01:35,439 --> 00:01:37,119 analogy you would give is that it is the 34 00:01:37,119 --> 00:01:40,000 digital backbone of Australia. 35 00:01:40,000 --> 00:01:41,759 And within such a critical organization 36 00:01:41,759 --> 00:01:43,759 such as this, as we have emerging 37 00:01:43,759 --> 00:01:47,759 technology like AI come about, we need a 38 00:01:47,759 --> 00:01:50,399 safe space within this organization to 39 00:01:50,399 --> 00:01:53,119 experiment. So that's what the incubator 40 00:01:53,119 --> 00:01:55,680 is. And in terms of the insights that 41 00:01:55,680 --> 00:01:57,759 I'll seek to share today, they have come 42 00:01:57,759 --> 00:02:00,960 from our experiment named Archimedes. Uh 43 00:02:00,960 --> 00:02:05,040 named after the individual who taught us 44 00:02:05,040 --> 00:02:08,640 about leverage, small input, high value 45 00:02:08,640 --> 00:02:10,239 output. 46 00:02:10,239 --> 00:02:13,280 Now, I'll touch on that very soon. 47 00:02:13,280 --> 00:02:15,040 But first, I would like you all to 48 00:02:15,040 --> 00:02:16,959 introspect and reflect about your 49 00:02:16,959 --> 00:02:19,680 experience over the past 11 months, 50 00:02:19,680 --> 00:02:22,080 roughly from October last year. What has 51 00:02:22,080 --> 00:02:24,959 your agentic AI journey been like? Has 52 00:02:24,959 --> 00:02:28,319 it been steady? Has it been a sore tooth 53 00:02:28,319 --> 00:02:30,239 where you've had some success followed 54 00:02:30,239 --> 00:02:32,080 by a bit of failure and then success 55 00:02:32,080 --> 00:02:36,080 again? or graph C. Have you had a 56 00:02:36,080 --> 00:02:37,920 tumultuous journey where your first 57 00:02:37,920 --> 00:02:40,560 prototype has looked amazing and you've 58 00:02:40,560 --> 00:02:42,800 struggled to get to that uh same height 59 00:02:42,800 --> 00:02:45,716 as that first week? Our team 60 00:02:45,716 --> 00:02:47,040 [sighs and gasps] within the incubator 61 00:02:47,040 --> 00:02:48,720 has had a combination of all of these 62 00:02:48,720 --> 00:02:50,800 different experiences. 63 00:02:50,800 --> 00:02:53,920 And what has felt 64 00:02:53,920 --> 00:02:56,879 that has blocked our progress with 65 00:02:56,879 --> 00:02:59,280 agentic development and agentic AI has 66 00:02:59,280 --> 00:03:01,680 been our expectations of what agents 67 00:03:01,680 --> 00:03:03,280 will do for us. 68 00:03:03,280 --> 00:03:05,680 versus the reality of what they're doing 69 00:03:05,680 --> 00:03:07,599 for us. 70 00:03:07,599 --> 00:03:10,080 So what we're what were we seeking from 71 00:03:10,080 --> 00:03:13,519 our agents? We were seeking competence. 72 00:03:13,519 --> 00:03:16,319 Our agents had a lot of capability, but 73 00:03:16,319 --> 00:03:19,599 they did not have competence. And as a 74 00:03:19,599 --> 00:03:20,720 team that's seeking to experiment with 75 00:03:20,720 --> 00:03:22,480 this within our organization to provide 76 00:03:22,480 --> 00:03:24,400 value, the question we're asking 77 00:03:24,400 --> 00:03:26,319 ourselves is we need to find out how to 78 00:03:26,319 --> 00:03:28,800 make them more competent. And a lot of 79 00:03:28,800 --> 00:03:31,840 our team are software engineers. And we 80 00:03:31,840 --> 00:03:32,959 thought we'd lean on the software 81 00:03:32,959 --> 00:03:34,640 engineering practice. What do software 82 00:03:34,640 --> 00:03:36,319 engineers do when we are faced with 83 00:03:36,319 --> 00:03:38,560 complexity and we are seeking to deliver 84 00:03:38,560 --> 00:03:40,159 value? 85 00:03:40,159 --> 00:03:42,799 We look to abstractions. Now here I 86 00:03:42,799 --> 00:03:44,799 don't have a an exhaustive list of 87 00:03:44,799 --> 00:03:46,400 traditional abstractions but it's just 88 00:03:46,400 --> 00:03:49,280 to set the example. We've got functions, 89 00:03:49,280 --> 00:03:52,400 classes, modules, services, all these 90 00:03:52,400 --> 00:03:55,280 different Lego blocks that have a name, 91 00:03:55,280 --> 00:03:57,599 a boundary, an input and output 92 00:03:57,599 --> 00:03:59,519 contract. And the cool thing about these 93 00:03:59,519 --> 00:04:02,480 abstractions is you can compose them 94 00:04:02,480 --> 00:04:04,480 together, build them up to deliver some 95 00:04:04,480 --> 00:04:06,560 value as a service. And we've already 96 00:04:06,560 --> 00:04:08,879 seen that in the AI lane, we do have 97 00:04:08,879 --> 00:04:11,040 abstractions, we do have names, we do 98 00:04:11,040 --> 00:04:13,040 have boundaries for certain areas of 99 00:04:13,040 --> 00:04:15,680 concern. And we've got input and output 100 00:04:15,680 --> 00:04:17,680 contracts, some in natural language, 101 00:04:17,680 --> 00:04:20,160 some in more deterministic interfaces. 102 00:04:20,160 --> 00:04:21,759 So we've got large language model we're 103 00:04:21,759 --> 00:04:23,199 all familiar with. We've got prompts as 104 00:04:23,199 --> 00:04:25,600 an abstraction tools and we've got an 105 00:04:25,600 --> 00:04:27,680 agent that is the composition of some of 106 00:04:27,680 --> 00:04:30,000 these abstractions together. But again, 107 00:04:30,000 --> 00:04:31,440 as we're going through our tumultuous 108 00:04:31,440 --> 00:04:33,520 journey that I described before, we were 109 00:04:33,520 --> 00:04:35,520 asking ourselves because this is the 110 00:04:35,520 --> 00:04:36,880 only language that we knew how to use at 111 00:04:36,880 --> 00:04:40,400 that time, is there an abstraction for 112 00:04:40,400 --> 00:04:43,040 agentic behavior where those sets of 113 00:04:43,040 --> 00:04:46,560 behaviors can then get us to a competent 114 00:04:46,560 --> 00:04:48,560 output that at least our organization 115 00:04:48,560 --> 00:04:51,680 would find valuable. Well, fast forward 116 00:04:51,680 --> 00:04:53,680 enter the abstraction for agentic 117 00:04:53,680 --> 00:04:56,560 behavior with a better brand than that. 118 00:04:56,560 --> 00:04:58,880 Agent skills. 119 00:04:58,880 --> 00:05:01,040 So, what are agent skills? The way I'd 120 00:05:01,040 --> 00:05:02,560 define them or describe them is that 121 00:05:02,560 --> 00:05:05,440 they are a modular package of 122 00:05:05,440 --> 00:05:07,199 instructions or procedures that allow 123 00:05:07,199 --> 00:05:10,720 you to optionally add in knowledge or 124 00:05:10,720 --> 00:05:13,039 references or deterministic scripts and 125 00:05:13,039 --> 00:05:16,000 tools. And the problem that they solve 126 00:05:16,000 --> 00:05:18,960 for us is that large language models are 127 00:05:18,960 --> 00:05:21,440 very capable. They have a lot of general 128 00:05:21,440 --> 00:05:23,680 knowledge, but they don't necessarily 129 00:05:23,680 --> 00:05:25,759 have our organization's knowledge. They 130 00:05:25,759 --> 00:05:28,000 don't have our organizations procedures 131 00:05:28,000 --> 00:05:31,680 or practices or decisions. So once we 132 00:05:31,680 --> 00:05:34,639 were able to observe u the value that 133 00:05:34,639 --> 00:05:37,440 anthropic provided the community with 134 00:05:37,440 --> 00:05:39,199 agent skills, we thought let's use these 135 00:05:39,199 --> 00:05:41,120 Lego blocks. Let's use these little 136 00:05:41,120 --> 00:05:43,199 modules, attach them to our agent 137 00:05:43,199 --> 00:05:45,360 harness that we're building, and see how 138 00:05:45,360 --> 00:05:47,759 much value we can get from this. Um, an 139 00:05:47,759 --> 00:05:50,400 extra benefit is that agent skills are 140 00:05:50,400 --> 00:05:53,440 an open standard. So, if you put a lot 141 00:05:53,440 --> 00:05:55,440 of energy and toil into creating an 142 00:05:55,440 --> 00:05:59,360 agent skill, you can get uh continuity 143 00:05:59,360 --> 00:06:00,560 of that skill across different 144 00:06:00,560 --> 00:06:01,759 harnesses, which is a really big 145 00:06:01,759 --> 00:06:04,759 positive. 146 00:06:04,839 --> 00:06:05,840 [sighs] 147 00:06:05,840 --> 00:06:08,639 Agent skill anatomy and integration. So 148 00:06:08,639 --> 00:06:10,160 how do you actually implement a skill? 149 00:06:10,160 --> 00:06:11,759 I've described that that it is a Lego 150 00:06:11,759 --> 00:06:14,479 block for example. It is essentially a 151 00:06:14,479 --> 00:06:16,720 folder. Uh the mandatory file that you 152 00:06:16,720 --> 00:06:19,120 need within this folder is the skill.md 153 00:06:19,120 --> 00:06:22,400 file. And in that particular file you'll 154 00:06:22,400 --> 00:06:24,960 have some YAML front matter which will 155 00:06:24,960 --> 00:06:26,639 define the name of your skill and 156 00:06:26,639 --> 00:06:28,720 importantly this description of that 157 00:06:28,720 --> 00:06:30,639 skill and in what instance the model 158 00:06:30,639 --> 00:06:32,479 should choose to use it for a particular 159 00:06:32,479 --> 00:06:34,080 task. 160 00:06:34,080 --> 00:06:36,880 And we can see also the effect of using 161 00:06:36,880 --> 00:06:38,720 skills on our context window which is a 162 00:06:38,720 --> 00:06:41,039 very important component of our agent 163 00:06:41,039 --> 00:06:42,800 and our model and how we're trying to 164 00:06:42,800 --> 00:06:45,520 deliver some value. And initially 165 00:06:45,520 --> 00:06:47,199 and the number of skills that we have 166 00:06:47,199 --> 00:06:49,360 available within our agent are exposed 167 00:06:49,360 --> 00:06:52,240 via uh skill catalog where only the name 168 00:06:52,240 --> 00:06:54,080 and descriptions are within the context 169 00:06:54,080 --> 00:06:56,400 window. And when the model decides it's 170 00:06:56,400 --> 00:06:58,639 worthy for a skill to be invoked or 171 00:06:58,639 --> 00:07:00,960 selected, will the skill MD files 172 00:07:00,960 --> 00:07:02,720 content be then stuffed into that 173 00:07:02,720 --> 00:07:05,360 window? And what's been really valuable 174 00:07:05,360 --> 00:07:06,960 about this type of abstraction and 175 00:07:06,960 --> 00:07:08,880 approach and the implementation that the 176 00:07:08,880 --> 00:07:10,479 standard recommends for harnesses to 177 00:07:10,479 --> 00:07:13,199 follow is that if you have added extra 178 00:07:13,199 --> 00:07:15,199 scripts for determinism, if you have 179 00:07:15,199 --> 00:07:16,960 added in references that are unique to 180 00:07:16,960 --> 00:07:19,840 your organization or your practice, um 181 00:07:19,840 --> 00:07:22,319 they can be progressively disclosed into 182 00:07:22,319 --> 00:07:24,880 the context window as opposed to just 183 00:07:24,880 --> 00:07:26,800 being stuffed straight in from the get- 184 00:07:26,800 --> 00:07:29,120 go. 185 00:07:29,120 --> 00:07:32,400 Now we asked ourselves where within our 186 00:07:32,400 --> 00:07:34,319 organization within our experiments can 187 00:07:34,319 --> 00:07:36,639 we seek to apply and test and verify 188 00:07:36,639 --> 00:07:39,280 agent skills would be of value. And we 189 00:07:39,280 --> 00:07:40,800 already had an experiment running which 190 00:07:40,800 --> 00:07:42,639 I introduced very briefly before named 191 00:07:42,639 --> 00:07:45,759 Archimedes which was our agent to 192 00:07:45,759 --> 00:07:47,680 support our enterprise architects in 193 00:07:47,680 --> 00:07:51,120 their role. Our experiment was to test 194 00:07:51,120 --> 00:07:54,160 could we have a natural language input 195 00:07:54,160 --> 00:07:57,120 business demand come in from our 196 00:07:57,120 --> 00:08:00,160 organization into an architect's queue 197 00:08:00,160 --> 00:08:01,680 and the architect would they would 198 00:08:01,680 --> 00:08:04,160 commonly judge whether or not that 199 00:08:04,160 --> 00:08:06,560 feature can land in our architecture or 200 00:08:06,560 --> 00:08:08,720 our solution and reason about whether 201 00:08:08,720 --> 00:08:11,840 it's appropriate to do 202 00:08:11,840 --> 00:08:14,240 that feature versus another feature and 203 00:08:14,240 --> 00:08:17,680 apply experience knowledge from um how 204 00:08:17,680 --> 00:08:19,840 these features been done in the past 205 00:08:19,840 --> 00:08:21,520 also against our current future 206 00:08:21,520 --> 00:08:23,280 architecture any decision records and 207 00:08:23,280 --> 00:08:25,280 principles and standards to then finally 208 00:08:25,280 --> 00:08:27,120 provide an output whether or not this 209 00:08:27,120 --> 00:08:29,599 feature will have an high impact or low 210 00:08:29,599 --> 00:08:31,599 impact on our architecture. It's really 211 00:08:31,599 --> 00:08:34,000 important within our organization to 212 00:08:34,000 --> 00:08:36,080 have a strategic architecture that can 213 00:08:36,080 --> 00:08:37,839 evolve over time and not have a lot of 214 00:08:37,839 --> 00:08:39,440 technical debt and the enterprise 215 00:08:39,440 --> 00:08:41,279 architect has to reason about a lot of 216 00:08:41,279 --> 00:08:43,519 different things. Our hypothesis was 217 00:08:43,519 --> 00:08:46,399 that an agent would be able to surface 218 00:08:46,399 --> 00:08:48,399 information when paired with an 219 00:08:48,399 --> 00:08:51,279 architect extremely quickly to help 220 00:08:51,279 --> 00:08:52,959 provide value to our organization in 221 00:08:52,959 --> 00:08:55,279 terms of speed. But then the quality is 222 00:08:55,279 --> 00:08:56,800 what we needed. We needed to ensure that 223 00:08:56,800 --> 00:08:58,800 our enterprise architect wasn't uh 224 00:08:58,800 --> 00:09:00,959 wasting time uh trolling through a first 225 00:09:00,959 --> 00:09:03,519 draft that is not of high quality. So 226 00:09:03,519 --> 00:09:05,360 the challenge was on us, the software 227 00:09:05,360 --> 00:09:06,720 engineering team, the agentic 228 00:09:06,720 --> 00:09:08,399 engineering team as we're evolving to 229 00:09:08,399 --> 00:09:10,480 that discipline. How can we make this 230 00:09:10,480 --> 00:09:13,040 agent provide initial drafts of these 231 00:09:13,040 --> 00:09:14,959 output artifacts which are essentially 232 00:09:14,959 --> 00:09:16,640 large documents that need to follow a 233 00:09:16,640 --> 00:09:18,720 particular format that would provide us 234 00:09:18,720 --> 00:09:21,720 value? 235 00:09:22,640 --> 00:09:24,480 Now what did it look like in terms of a 236 00:09:24,480 --> 00:09:26,000 high level implementation? We used 237 00:09:26,000 --> 00:09:28,720 Langraph as our own agent harness. We 238 00:09:28,720 --> 00:09:32,320 had a research plan, execute, replan 239 00:09:32,320 --> 00:09:35,279 node structure. And our methodology at 240 00:09:35,279 --> 00:09:36,800 that time before agent skills were 241 00:09:36,800 --> 00:09:38,640 introduced was to provide what we call 242 00:09:38,640 --> 00:09:41,360 golden examples of impact assessments, 243 00:09:41,360 --> 00:09:43,519 golden examples of decision records, 244 00:09:43,519 --> 00:09:46,000 golden examples of solutions on a page. 245 00:09:46,000 --> 00:09:50,000 And our hope was that the model would 246 00:09:50,000 --> 00:09:52,480 have sufficient enough inference to 247 00:09:52,480 --> 00:09:54,480 determine given a question that is asked 248 00:09:54,480 --> 00:09:57,519 of an architect or an architect asked 249 00:09:57,519 --> 00:10:00,000 the agent and the golden examples it has 250 00:10:00,000 --> 00:10:01,920 available, it would be able to infer the 251 00:10:01,920 --> 00:10:03,839 appropriate reasoning steps to get there 252 00:10:03,839 --> 00:10:05,680 whilst also retrieving data from the 253 00:10:05,680 --> 00:10:08,000 appropriate locations where necessary. 254 00:10:08,000 --> 00:10:09,360 Um, you might [snorts] also notice that 255 00:10:09,360 --> 00:10:11,680 there's uh traces here, which was 256 00:10:11,680 --> 00:10:13,839 extremely important because we needed to 257 00:10:13,839 --> 00:10:15,839 observe how our architects were going to 258 00:10:15,839 --> 00:10:18,480 be using our agent to ensure that we had 259 00:10:18,480 --> 00:10:19,839 visibility over the types of questions 260 00:10:19,839 --> 00:10:21,200 they were asking in particular when 261 00:10:21,200 --> 00:10:23,839 things were going wrong. So, what issues 262 00:10:23,839 --> 00:10:25,600 did we experience without agent skills 263 00:10:25,600 --> 00:10:28,320 in this particular harness? 264 00:10:28,320 --> 00:10:30,399 I shared with you our golden example, 265 00:10:30,399 --> 00:10:33,519 our golden example example, and that was 266 00:10:33,519 --> 00:10:36,399 unreliable. So we noticed that the 267 00:10:36,399 --> 00:10:38,800 behaviors of our agent was different for 268 00:10:38,800 --> 00:10:41,600 every unique run that the architects 269 00:10:41,600 --> 00:10:44,320 were asking questions of. Our context 270 00:10:44,320 --> 00:10:45,760 window had a lot of pressure. You 271 00:10:45,760 --> 00:10:47,440 noticed perhaps in that previous diagram 272 00:10:47,440 --> 00:10:49,839 that we had MCP reference to have access 273 00:10:49,839 --> 00:10:52,000 to our data stores and all those tool 274 00:10:52,000 --> 00:10:53,360 definitions and resource definitions 275 00:10:53,360 --> 00:10:55,279 were taking up a lot of context that 276 00:10:55,279 --> 00:10:57,839 wasn't necessary. And we had a number of 277 00:10:57,839 --> 00:10:59,200 examples where we thought let's just 278 00:10:59,200 --> 00:11:01,279 throw the whole kitchen sink at this and 279 00:11:01,279 --> 00:11:03,200 provide as much information to our agent 280 00:11:03,200 --> 00:11:04,880 as possible for it to be able to reason 281 00:11:04,880 --> 00:11:07,120 about how to get to the target output. 282 00:11:07,120 --> 00:11:08,720 We also had a lot of challenges with 283 00:11:08,720 --> 00:11:11,440 similarity search surfacing information 284 00:11:11,440 --> 00:11:14,720 from our uh architectural repositories 285 00:11:14,720 --> 00:11:16,800 that were somewhat similar vaguely 286 00:11:16,800 --> 00:11:18,720 relevant but not the right document to 287 00:11:18,720 --> 00:11:20,640 actually give us the first draft that 288 00:11:20,640 --> 00:11:23,040 our architects needed on a solution on a 289 00:11:23,040 --> 00:11:25,847 page. [snorts] 290 00:11:26,640 --> 00:11:28,079 What did we learn when we implemented 291 00:11:28,079 --> 00:11:30,160 agent skills? 292 00:11:30,160 --> 00:11:32,240 We learned immediately how powerful it 293 00:11:32,240 --> 00:11:36,079 is for a harness to be in a loop and to 294 00:11:36,079 --> 00:11:39,839 have skills act to constrain 295 00:11:39,839 --> 00:11:41,519 that orchestration loop and the 296 00:11:41,519 --> 00:11:43,839 reasoning loop of our agent. So we got 297 00:11:43,839 --> 00:11:46,720 to simplify our node structure to just a 298 00:11:46,720 --> 00:11:48,640 couple of nodes. We still had the same 299 00:11:48,640 --> 00:11:53,120 capabilities stuffed into our agent. we 300 00:11:53,120 --> 00:11:56,240 just had more competence. So that's how 301 00:11:56,240 --> 00:11:59,200 the the harness changed. 302 00:11:59,200 --> 00:12:01,760 But in terms of the good, 303 00:12:01,760 --> 00:12:04,079 we noticed immediately that we had a 304 00:12:04,079 --> 00:12:08,079 reason to the method to the madness. We 305 00:12:08,079 --> 00:12:09,760 had the previous sprawl within our 306 00:12:09,760 --> 00:12:12,000 harness reduced. We had more of a 307 00:12:12,000 --> 00:12:13,440 modular approach. And this is really 308 00:12:13,440 --> 00:12:15,360 important as we're trying to scale out 309 00:12:15,360 --> 00:12:18,000 an agent at runtime to be able to reason 310 00:12:18,000 --> 00:12:19,600 about how to make changes, which is the 311 00:12:19,600 --> 00:12:21,440 second point. Once you have a modular 312 00:12:21,440 --> 00:12:22,800 design and you know where to make 313 00:12:22,800 --> 00:12:24,560 changes, it's much faster in terms of 314 00:12:24,560 --> 00:12:26,480 your feedback loop to add further 315 00:12:26,480 --> 00:12:28,240 constraints to your skill to help 316 00:12:28,240 --> 00:12:30,560 confine it. Uh extra deterministic 317 00:12:30,560 --> 00:12:32,240 scripts once you learn that your scripts 318 00:12:32,240 --> 00:12:34,720 are having issues and adding extra 319 00:12:34,720 --> 00:12:37,040 gotcha sections again to en encourage 320 00:12:37,040 --> 00:12:40,079 and steer the agent and the harness to 321 00:12:40,079 --> 00:12:42,880 deliver an outcome. And another benefit 322 00:12:42,880 --> 00:12:44,880 of the agent skills open standard and 323 00:12:44,880 --> 00:12:47,680 our harness abiding by it was that our 324 00:12:47,680 --> 00:12:50,320 context was immediately lean and only 325 00:12:50,320 --> 00:12:53,760 progressively would it gain more 326 00:12:53,760 --> 00:12:55,920 information where the model decided to 327 00:12:55,920 --> 00:12:57,839 do so. So it had plenty more space to 328 00:12:57,839 --> 00:13:00,823 reason. [sighs] 329 00:13:01,279 --> 00:13:03,040 Another thing that was a benefit and the 330 00:13:03,040 --> 00:13:05,120 good from us implementing agent skills 331 00:13:05,120 --> 00:13:08,399 was the standard encouraging and 332 00:13:08,399 --> 00:13:10,560 proposing examples of moving towards a 333 00:13:10,560 --> 00:13:13,440 procedural how versus our previous 334 00:13:13,440 --> 00:13:14,880 approach before agent skills were 335 00:13:14,880 --> 00:13:16,639 released where we were taking more of an 336 00:13:16,639 --> 00:13:18,079 inference approach and putting a lot of 337 00:13:18,079 --> 00:13:20,959 pressure on the model to make decisions. 338 00:13:20,959 --> 00:13:22,800 And the benefit of going for a 339 00:13:22,800 --> 00:13:24,480 procedural how approach and within our 340 00:13:24,480 --> 00:13:26,959 organization having domain experts is 341 00:13:26,959 --> 00:13:30,480 that skill creation is intuitive. It's 342 00:13:30,480 --> 00:13:32,320 natural language. If you've already 343 00:13:32,320 --> 00:13:34,320 written a procedure, if you've already 344 00:13:34,320 --> 00:13:35,920 written a playbook before, you 345 00:13:35,920 --> 00:13:38,880 essentially know what a skill encodes in 346 00:13:38,880 --> 00:13:40,560 that markdown file. It encodes a 347 00:13:40,560 --> 00:13:42,639 procedure. It encodes references to 348 00:13:42,639 --> 00:13:43,999 further information. 349 00:13:43,999 --> 00:13:44,079 [snorts] 350 00:13:44,079 --> 00:13:45,680 Um the challenge we had with our 351 00:13:45,680 --> 00:13:48,240 semantic search providing only vaguely 352 00:13:48,240 --> 00:13:51,440 relevant documentation was solved via 353 00:13:51,440 --> 00:13:52,959 one mechanism here which we called our 354 00:13:52,959 --> 00:13:55,440 expert search skill which was to simply 355 00:13:55,440 --> 00:13:57,839 ask our enterprise architects how do you 356 00:13:57,839 --> 00:14:00,000 search through our data stores? Do you 357 00:14:00,000 --> 00:14:01,680 have a hierarchical approach in which 358 00:14:01,680 --> 00:14:04,560 you decide to prefer some documents over 359 00:14:04,560 --> 00:14:06,399 others? And if that's the case what are 360 00:14:06,399 --> 00:14:07,920 your rules? Can you please share them in 361 00:14:07,920 --> 00:14:11,279 natural language? And we had amazing 362 00:14:11,279 --> 00:14:14,129 success with our expert search skill. 363 00:14:14,129 --> 00:14:15,680 [snorts] Now, here's the point I want to 364 00:14:15,680 --> 00:14:17,199 share with you all. Complexity still 365 00:14:17,199 --> 00:14:20,560 existed. Complexity still exists, but 366 00:14:20,560 --> 00:14:23,519 it's managed by abstractions. 367 00:14:23,519 --> 00:14:24,720 It's almost like having that big 368 00:14:24,720 --> 00:14:26,079 elephant and just trying to reduce it 369 00:14:26,079 --> 00:14:28,240 into smaller chunks. 370 00:14:28,240 --> 00:14:30,399 Um, now in terms of challenges we faced, 371 00:14:30,399 --> 00:14:31,760 this is the the section of the talk 372 00:14:31,760 --> 00:14:34,160 where I want to share the challenges we 373 00:14:34,160 --> 00:14:35,920 had and hopefully they'll be insightful 374 00:14:35,920 --> 00:14:37,680 for you if you have similar challenges. 375 00:14:37,680 --> 00:14:39,839 And I've split them up into invocation 376 00:14:39,839 --> 00:14:41,519 challenges, invocation meaning 377 00:14:41,519 --> 00:14:43,920 triggering the skill, and structural 378 00:14:43,920 --> 00:14:45,920 challenges. So challenges where we 379 00:14:45,920 --> 00:14:47,440 eventually had to change the content 380 00:14:47,440 --> 00:14:50,000 structures of our skill. Here's an 381 00:14:50,000 --> 00:14:51,839 example. It's very common one. You put 382 00:14:51,839 --> 00:14:53,279 all this energy and effort into creating 383 00:14:53,279 --> 00:14:55,440 a skill and you're hoping that your 384 00:14:55,440 --> 00:14:57,920 agent and the model within it will 385 00:14:57,920 --> 00:15:01,290 select your skill to be invoked. 386 00:15:01,290 --> 00:15:02,079 [sighs] 387 00:15:02,079 --> 00:15:04,079 It was a challenge, [laughter] but we 388 00:15:04,079 --> 00:15:05,920 eventually had some sort of framework to 389 00:15:05,920 --> 00:15:07,839 follow and that was to within the 390 00:15:07,839 --> 00:15:09,600 description that very important metadata 391 00:15:09,600 --> 00:15:12,399 field make it clear when should your 392 00:15:12,399 --> 00:15:14,959 model and agent use this particular 393 00:15:14,959 --> 00:15:17,760 skill but also importantly name the 394 00:15:17,760 --> 00:15:21,440 artifacts that you expect from that 395 00:15:21,440 --> 00:15:23,519 skill because it may be common for your 396 00:15:23,519 --> 00:15:27,839 users to ask for a impact assessment as 397 00:15:27,839 --> 00:15:29,199 an artifact. So you want to make sure 398 00:15:29,199 --> 00:15:30,880 that that's within the description. But 399 00:15:30,880 --> 00:15:33,120 then you also want to name phrasings 400 00:15:33,120 --> 00:15:36,160 that are common for your user group uh 401 00:15:36,160 --> 00:15:37,839 for that particular skill. So in this 402 00:15:37,839 --> 00:15:39,440 case there is some colloquial language 403 00:15:39,440 --> 00:15:41,680 or tribal knowledge language like a BRD 404 00:15:41,680 --> 00:15:43,440 that I don't expect anyone else here to 405 00:15:43,440 --> 00:15:45,120 understand but it's important for at 406 00:15:45,120 --> 00:15:46,800 least within the skill for the model to 407 00:15:46,800 --> 00:15:49,360 see that type of information. 408 00:15:49,360 --> 00:15:50,959 Here's the opposite scenario where you 409 00:15:50,959 --> 00:15:53,120 have invocation collisions. You have a 410 00:15:53,120 --> 00:15:55,680 skill, execute, fantastic, but 411 00:15:55,680 --> 00:15:57,519 unfortunately 412 00:15:57,519 --> 00:15:59,040 the descriptions you've written for one 413 00:15:59,040 --> 00:16:01,600 skill which might have been fantastic 414 00:16:01,600 --> 00:16:03,360 could be ruined by the description you 415 00:16:03,360 --> 00:16:04,959 have in another skill that is less 416 00:16:04,959 --> 00:16:06,720 fantastic. 417 00:16:06,720 --> 00:16:09,360 So in our particular scenario, which may 418 00:16:09,360 --> 00:16:11,040 be a different context for most people 419 00:16:11,040 --> 00:16:12,800 here where they're maybe developing a 420 00:16:12,800 --> 00:16:15,120 coding agent versus a runtime agent 421 00:16:15,120 --> 00:16:17,360 where it's running and users are seeking 422 00:16:17,360 --> 00:16:19,759 to have value invoked. 423 00:16:19,759 --> 00:16:22,399 We as software engineers controlling our 424 00:16:22,399 --> 00:16:24,560 agent and the skills that we're putting 425 00:16:24,560 --> 00:16:26,079 into our agent need to be careful about 426 00:16:26,079 --> 00:16:28,079 the descriptions and potential clashes 427 00:16:28,079 --> 00:16:29,839 that those descriptions have within our 428 00:16:29,839 --> 00:16:32,240 product. So in this particular example, 429 00:16:32,240 --> 00:16:34,079 using the word impact assessment in our 430 00:16:34,079 --> 00:16:36,880 security assessment was the the wrong 431 00:16:36,880 --> 00:16:39,519 idea. So we essentially resolved that by 432 00:16:39,519 --> 00:16:42,720 ensuring um separation of concerns and 433 00:16:42,720 --> 00:16:45,519 differentiating our descriptions. 434 00:16:45,519 --> 00:16:47,040 Now, the last thing I want to mention on 435 00:16:47,040 --> 00:16:48,720 invocation 436 00:16:48,720 --> 00:16:52,399 is that we within our enterprise agent 437 00:16:52,399 --> 00:16:54,480 needed to make an invocation decision. 438 00:16:54,480 --> 00:16:56,160 There's a number of people in the 439 00:16:56,160 --> 00:16:57,199 audience right now that might be 440 00:16:57,199 --> 00:16:58,480 thinking, well, you've talked a lot 441 00:16:58,480 --> 00:17:00,160 about model invocation. What about user 442 00:17:00,160 --> 00:17:02,720 invocation via slash commands? So, by 443 00:17:02,720 --> 00:17:04,480 default, agent harnesses that comply to 444 00:17:04,480 --> 00:17:08,079 the standard will supply slash skill 445 00:17:08,079 --> 00:17:11,120 user invocation and naturally invocation 446 00:17:11,120 --> 00:17:14,559 of skills from your context. [snorts] 447 00:17:14,559 --> 00:17:16,640 You can make a decision to disable model 448 00:17:16,640 --> 00:17:18,799 invocation by setting a metadata flag. 449 00:17:18,799 --> 00:17:20,640 Um, not every harness has this very 450 00:17:20,640 --> 00:17:24,240 precise metadata key, but essentially 451 00:17:24,240 --> 00:17:26,160 there are harnesses that allow you to 452 00:17:26,160 --> 00:17:28,400 disable the invocation from a model, 453 00:17:28,400 --> 00:17:30,240 which would imply and put burden on your 454 00:17:30,240 --> 00:17:32,559 users to already be aware of what skill 455 00:17:32,559 --> 00:17:35,520 they wish to execute and expect them to 456 00:17:35,520 --> 00:17:37,837 do that whenever they wish to. The 457 00:17:37,837 --> 00:17:39,760 [snorts] benefit of enabling or pardon 458 00:17:39,760 --> 00:17:41,840 me disabling model invocation is that it 459 00:17:41,840 --> 00:17:43,679 saves context. you no longer have an 460 00:17:43,679 --> 00:17:45,360 array of skills that you've disabled 461 00:17:45,360 --> 00:17:46,880 from filling up the context window, 462 00:17:46,880 --> 00:17:49,520 which could be a value for you. For our 463 00:17:49,520 --> 00:17:51,039 scenario, 464 00:17:51,039 --> 00:17:53,280 we decided to create orchestration 465 00:17:53,280 --> 00:17:56,160 skills, which would be a skill, for 466 00:17:56,160 --> 00:17:58,559 example, skill A that contained a 467 00:17:58,559 --> 00:18:00,080 procedure, and this is a very simple 468 00:18:00,080 --> 00:18:02,880 example where within that procedure 469 00:18:02,880 --> 00:18:05,200 would be call outs to other skills in 470 00:18:05,200 --> 00:18:06,799 sequence. So, for example, skill B and 471 00:18:06,799 --> 00:18:09,280 skill C. So we could support either the 472 00:18:09,280 --> 00:18:11,919 model invoking skill A or we could 473 00:18:11,919 --> 00:18:14,480 support a user invoking skill A and then 474 00:18:14,480 --> 00:18:16,640 from that point onwards it would all be 475 00:18:16,640 --> 00:18:18,720 model invoked 476 00:18:18,720 --> 00:18:21,760 uh skills. So if [snorts] we decided as 477 00:18:21,760 --> 00:18:24,480 a group to have disabled model 478 00:18:24,480 --> 00:18:27,840 invocation on skill B and skill C 479 00:18:27,840 --> 00:18:30,799 despite an enterprise architect invoking 480 00:18:30,799 --> 00:18:33,520 skill A expecting it to orchestrate from 481 00:18:33,520 --> 00:18:36,320 there onwards the harness would reject 482 00:18:36,320 --> 00:18:38,799 invocation of skill B and C because it's 483 00:18:38,799 --> 00:18:41,520 now in the models domain. So I just want 484 00:18:41,520 --> 00:18:43,919 to share that we settled on uh the 485 00:18:43,919 --> 00:18:47,600 orchestration user and model approach 486 00:18:47,600 --> 00:18:50,080 in terms of structural challenges and 487 00:18:50,080 --> 00:18:51,520 things that required us to change the 488 00:18:51,520 --> 00:18:54,000 structure of our skill. Enterprise 489 00:18:54,000 --> 00:18:55,440 architecture is a big beast and there's 490 00:18:55,440 --> 00:18:58,000 a lot of reference files necessary and 491 00:18:58,000 --> 00:19:00,880 in our earlier implementations we had a 492 00:19:00,880 --> 00:19:03,120 lot of architecture principles stuffed 493 00:19:03,120 --> 00:19:06,080 into our main skill MD file and we 494 00:19:06,080 --> 00:19:10,080 noticed that we had dilution of value of 495 00:19:10,080 --> 00:19:12,799 our instructions where the instructions 496 00:19:12,799 --> 00:19:14,559 or principles in the middle were being 497 00:19:14,559 --> 00:19:16,720 ignored on particular enterprise domains 498 00:19:16,720 --> 00:19:20,000 that we needed to be honored. We 499 00:19:20,000 --> 00:19:22,880 followed the recommendation to have 500 00:19:22,880 --> 00:19:26,080 references for different principles that 501 00:19:26,080 --> 00:19:27,280 land in different domains in our 502 00:19:27,280 --> 00:19:30,000 organization instead being in separate 503 00:19:30,000 --> 00:19:31,919 reference file locations that's still 504 00:19:31,919 --> 00:19:34,080 within your skill folder. And the 505 00:19:34,080 --> 00:19:35,760 benefit of that was it essentially 506 00:19:35,760 --> 00:19:38,640 steered and encouraged the agent harness 507 00:19:38,640 --> 00:19:40,960 and the model to have more reasoning 508 00:19:40,960 --> 00:19:42,960 steps where each individual reasoning 509 00:19:42,960 --> 00:19:45,760 step was focusing on that particular 510 00:19:45,760 --> 00:19:48,080 domain. um as opposed to the first 511 00:19:48,080 --> 00:19:49,600 instance where you have potentially one 512 00:19:49,600 --> 00:19:51,440 reasoning step that's trying to do too 513 00:19:51,440 --> 00:19:54,080 much with too many instructions. So this 514 00:19:54,080 --> 00:19:56,000 is um something that perhaps you could 515 00:19:56,000 --> 00:19:59,600 uh take home as a thought. I mentioned 516 00:19:59,600 --> 00:20:03,600 our expert search skill and how we asked 517 00:20:03,600 --> 00:20:05,360 them to share with us precisely their 518 00:20:05,360 --> 00:20:07,360 procedure of going through our different 519 00:20:07,360 --> 00:20:09,520 data stores and how they get value and 520 00:20:09,520 --> 00:20:10,640 the information they need to make 521 00:20:10,640 --> 00:20:13,600 decisions. And initially when we invoked 522 00:20:13,600 --> 00:20:15,760 that skill, it was fantastic at giving 523 00:20:15,760 --> 00:20:17,600 us the information we needed, but it 524 00:20:17,600 --> 00:20:19,919 filled the context window significantly 525 00:20:19,919 --> 00:20:22,080 on the main context. And what we 526 00:20:22,080 --> 00:20:23,919 realized is that it'd be great if we 527 00:20:23,919 --> 00:20:25,679 could have this delegated to a sub 528 00:20:25,679 --> 00:20:28,159 agent. Um, which some harnesses allow 529 00:20:28,159 --> 00:20:31,039 you to set a metadata flag in their 530 00:20:31,039 --> 00:20:32,960 metadata that says, please run this on a 531 00:20:32,960 --> 00:20:34,720 sub aent, but not all harnesses support 532 00:20:34,720 --> 00:20:36,880 that. But what you can do in natural 533 00:20:36,880 --> 00:20:39,919 language is within your skillmd file 534 00:20:39,919 --> 00:20:42,159 steer agent orchestration. At the end of 535 00:20:42,159 --> 00:20:44,320 the day, the agent/model makes the 536 00:20:44,320 --> 00:20:46,480 decision on how it orchestrates things. 537 00:20:46,480 --> 00:20:48,400 But part of the skill's purpose is to 538 00:20:48,400 --> 00:20:50,960 encourage particular behavior. And the 539 00:20:50,960 --> 00:20:54,080 behavior we were seeking was sub agent 540 00:20:54,080 --> 00:20:56,720 execution and the final response 541 00:20:56,720 --> 00:21:00,880 surfacing back into the main context. 542 00:21:00,880 --> 00:21:02,320 I'm sure a lot of you have experienced 543 00:21:02,320 --> 00:21:04,320 this as well that are playing with agent 544 00:21:04,320 --> 00:21:07,280 skills. You'll set a task into your 545 00:21:07,280 --> 00:21:08,960 harness. The harness will decide to 546 00:21:08,960 --> 00:21:10,880 execute your skill. The skill has a 547 00:21:10,880 --> 00:21:12,159 particular goal it's seeking to 548 00:21:12,159 --> 00:21:15,120 implement or get reach, but it goes off 549 00:21:15,120 --> 00:21:16,559 track. It might even reach your goal but 550 00:21:16,559 --> 00:21:18,799 decide to do even more. Um, our target 551 00:21:18,799 --> 00:21:21,120 state is for it to exit early or to 552 00:21:21,120 --> 00:21:23,039 reach its goal successfully and move on 553 00:21:23,039 --> 00:21:25,520 and not do anything more. So our team 554 00:21:25,520 --> 00:21:27,200 eventually moved towards the practice of 555 00:21:27,200 --> 00:21:30,080 ensuring we have early exits, 556 00:21:30,080 --> 00:21:32,000 a a section of gotchas where we can 557 00:21:32,000 --> 00:21:34,799 continue to append more information to 558 00:21:34,799 --> 00:21:37,679 and a done when clause and you start to 559 00:21:37,679 --> 00:21:39,200 see that a lot of this begins to look 560 00:21:39,200 --> 00:21:41,600 like the shape of deterministic 561 00:21:41,600 --> 00:21:43,520 traditional software engineering but 562 00:21:43,520 --> 00:21:45,840 more into natural language uh type 563 00:21:45,840 --> 00:21:49,039 constraint. The last challenge I'll 564 00:21:49,039 --> 00:21:52,400 finish on is 565 00:21:52,400 --> 00:21:54,799 procedural reinvention. So, we have live 566 00:21:54,799 --> 00:21:56,799 traces from our agent that's sharing 567 00:21:56,799 --> 00:21:59,200 with us all of the different decisions 568 00:21:59,200 --> 00:22:01,679 that the agent's making. And for some 569 00:22:01,679 --> 00:22:03,760 instances where our agent and the skill 570 00:22:03,760 --> 00:22:06,480 demands that it runs code, it starts to 571 00:22:06,480 --> 00:22:09,039 invent procedures and execute those 572 00:22:09,039 --> 00:22:12,559 scripts. And unfortunately for us, given 573 00:22:12,559 --> 00:22:14,400 as a business, we've got cost 574 00:22:14,400 --> 00:22:16,799 constraints. Uh if we can move things 575 00:22:16,799 --> 00:22:18,880 into determinism, we should absolutely 576 00:22:18,880 --> 00:22:20,000 do that. That would be the 577 00:22:20,000 --> 00:22:21,760 recommendation here for not only 578 00:22:21,760 --> 00:22:23,760 yourselves but our team. Create 579 00:22:23,760 --> 00:22:25,440 deterministic scripts based off prior 580 00:22:25,440 --> 00:22:29,520 history and put them into your skill. 581 00:22:29,520 --> 00:22:31,200 So we had a lot of challenges. Hope you 582 00:22:31,200 --> 00:22:33,280 get some value from sharing me sharing 583 00:22:33,280 --> 00:22:35,520 that. Um so what surprised us in this 584 00:22:35,520 --> 00:22:39,039 agentic AI journey that we embarked on. 585 00:22:39,039 --> 00:22:40,640 We were predominantly software engineers 586 00:22:40,640 --> 00:22:44,000 initially. Then AI has arrived and been 587 00:22:44,000 --> 00:22:46,159 moving over to a nondeterministic world. 588 00:22:46,159 --> 00:22:48,720 We've had to be okay as engineers. 589 00:22:48,720 --> 00:22:51,760 widening our net of what's acceptable 590 00:22:51,760 --> 00:22:54,559 from the outcome from an agent. And 591 00:22:54,559 --> 00:22:57,200 along the same train of thought is we've 592 00:22:57,200 --> 00:23:00,000 now become constraint engineers. We have 593 00:23:00,000 --> 00:23:01,360 natural language coming in. We've got 594 00:23:01,360 --> 00:23:03,120 this abstraction available to us which 595 00:23:03,120 --> 00:23:05,120 is an agent skill. What are we seeking 596 00:23:05,120 --> 00:23:06,640 to do with an agent skill? Of course, 597 00:23:06,640 --> 00:23:08,159 provide some sort of example of a 598 00:23:08,159 --> 00:23:09,919 procedure with reference files so on and 599 00:23:09,919 --> 00:23:11,760 so forth. 600 00:23:11,760 --> 00:23:15,280 But we're seeking to constrain to to as 601 00:23:15,280 --> 00:23:17,840 best ability as we can how that skill 602 00:23:17,840 --> 00:23:19,760 will operate so that it doesn't go off 603 00:23:19,760 --> 00:23:21,840 the off the rails. And again, the 604 00:23:21,840 --> 00:23:24,240 context here for our organization was to 605 00:23:24,240 --> 00:23:27,440 ensure we had a higher reliability of uh 606 00:23:27,440 --> 00:23:30,240 output from our agents. So given that 607 00:23:30,240 --> 00:23:32,400 context, we obviously were focusing more 608 00:23:32,400 --> 00:23:34,480 on constraining skills as opposed to 609 00:23:34,480 --> 00:23:37,200 letting them run wild. Another thing 610 00:23:37,200 --> 00:23:38,480 that surprised us which was really 611 00:23:38,480 --> 00:23:42,000 positive was the composition of natural 612 00:23:42,000 --> 00:23:45,280 language as an input to deliver some 613 00:23:45,280 --> 00:23:47,200 sort of capability while still 614 00:23:47,200 --> 00:23:50,559 maintaining the ability to compose 615 00:23:50,559 --> 00:23:53,280 scripts through interface. So natural 616 00:23:53,280 --> 00:23:54,720 language would kick off a skill that 617 00:23:54,720 --> 00:23:56,320 would follow through its procedure and 618 00:23:56,320 --> 00:23:58,480 the model would decide based off the 619 00:23:58,480 --> 00:23:59,840 information in the skill how it would 620 00:23:59,840 --> 00:24:01,840 proceed onwards. But we still maintained 621 00:24:01,840 --> 00:24:04,240 our ability to also compose not only 622 00:24:04,240 --> 00:24:05,679 through context natural language but 623 00:24:05,679 --> 00:24:07,919 compose through interface which is the 624 00:24:07,919 --> 00:24:09,520 traditional software engineering input 625 00:24:09,520 --> 00:24:11,600 output contracts. Our [snorts] 626 00:24:11,600 --> 00:24:14,240 mental model changed from prompting to 627 00:24:14,240 --> 00:24:17,919 explore and then moving to to skill 628 00:24:17,919 --> 00:24:19,679 creation when it was clear that value 629 00:24:19,679 --> 00:24:22,559 was worth retaining and as you've seen 630 00:24:22,559 --> 00:24:25,120 me mention before moving deterministic 631 00:24:25,120 --> 00:24:28,799 things to to scripts to avoid procedural 632 00:24:28,799 --> 00:24:30,480 reinvention. 633 00:24:30,480 --> 00:24:33,600 I've also got shown there traces. 634 00:24:33,600 --> 00:24:35,200 I've mentioned traces a couple of times. 635 00:24:35,200 --> 00:24:36,720 It gives us visibility into what's going 636 00:24:36,720 --> 00:24:39,120 on in our agent and empowers us to again 637 00:24:39,120 --> 00:24:41,360 increase our reliability. We had an 638 00:24:41,360 --> 00:24:45,120 ability to observe the shape of our uh 639 00:24:45,120 --> 00:24:47,679 traces and request that our architects 640 00:24:47,679 --> 00:24:49,679 were continuously asking which gave us 641 00:24:49,679 --> 00:24:51,760 an opportunity to skill spot or skill 642 00:24:51,760 --> 00:24:54,240 mine and prevent or present skilled 643 00:24:54,240 --> 00:24:56,720 candidates. 644 00:24:56,720 --> 00:24:58,320 I've talked a lot about challenges. I've 645 00:24:58,320 --> 00:24:59,520 talked a little bit there about skill 646 00:24:59,520 --> 00:25:02,000 spotting. I've talked about changing our 647 00:25:02,000 --> 00:25:04,720 skills. But if you change your skills 648 00:25:04,720 --> 00:25:07,679 and then put it in front of your user 649 00:25:07,679 --> 00:25:10,240 group without evaluating or testing 650 00:25:10,240 --> 00:25:11,840 them, you're essentially experimenting 651 00:25:11,840 --> 00:25:13,679 on your user group, which is not a good 652 00:25:13,679 --> 00:25:16,080 idea. So for us to be confident that any 653 00:25:16,080 --> 00:25:18,640 skills that we're modifying will provide 654 00:25:18,640 --> 00:25:21,520 value, we need evaluations. 655 00:25:21,520 --> 00:25:22,799 And if there's another thing that you'll 656 00:25:22,799 --> 00:25:24,720 take away from this talk, please is is 657 00:25:24,720 --> 00:25:26,400 if you're serious about improving the 658 00:25:26,400 --> 00:25:28,880 reliability of your agent and you're 659 00:25:28,880 --> 00:25:31,200 putting agents in front of users, please 660 00:25:31,200 --> 00:25:32,960 evaluate your agents. But in this 661 00:25:32,960 --> 00:25:34,640 instance, since I'm focusing on skills, 662 00:25:34,640 --> 00:25:37,120 evaluate your agent that is invoking 663 00:25:37,120 --> 00:25:39,840 your skills. You want realistic input 664 00:25:39,840 --> 00:25:41,679 prompts. Um, so in our instance, we had 665 00:25:41,679 --> 00:25:43,039 realistic input prompts from our 666 00:25:43,039 --> 00:25:45,520 enterprise architects. And we also 667 00:25:45,520 --> 00:25:47,200 defined with help from our enterprise 668 00:25:47,200 --> 00:25:49,120 architects how they would evaluate the 669 00:25:49,120 --> 00:25:51,840 output from an agent in both a 670 00:25:51,840 --> 00:25:53,679 deterministic way with for example 671 00:25:53,679 --> 00:25:55,360 formatting checks but also a 672 00:25:55,360 --> 00:25:57,520 nondeterministic way similar to a 673 00:25:57,520 --> 00:26:00,159 teacher grading homework um a marking 674 00:26:00,159 --> 00:26:01,600 guide for a rubric of how good that 675 00:26:01,600 --> 00:26:04,080 output was. Once we had those two things 676 00:26:04,080 --> 00:26:06,240 the prompts and the evaluation criteria 677 00:26:06,240 --> 00:26:09,360 would then move towards uh running the 678 00:26:09,360 --> 00:26:12,480 agent on at least two different arms. 679 00:26:12,480 --> 00:26:14,799 Um, one would be the previous skill that 680 00:26:14,799 --> 00:26:17,279 we considered a baseline skill and our 681 00:26:17,279 --> 00:26:19,120 version two skill which is after we made 682 00:26:19,120 --> 00:26:20,880 some changes and only in the instance 683 00:26:20,880 --> 00:26:23,520 that we had seen an improvement into the 684 00:26:23,520 --> 00:26:26,480 output where we can leverage an LLM as a 685 00:26:26,480 --> 00:26:29,279 judge on a separate lane would we seek 686 00:26:29,279 --> 00:26:31,600 to promote that particular skills 687 00:26:31,600 --> 00:26:34,559 version to our user group. Now it's 688 00:26:34,559 --> 00:26:36,720 Pyon. [laughter] I need to share some 689 00:26:36,720 --> 00:26:38,720 Python of some description. This is a 690 00:26:38,720 --> 00:26:40,240 very simple example for those people 691 00:26:40,240 --> 00:26:41,919 that want to get started. And my main 692 00:26:41,919 --> 00:26:44,159 goal here is to encourage evaluation. 693 00:26:44,159 --> 00:26:47,120 Our team spent a lot of energy and time 694 00:26:47,120 --> 00:26:49,919 using deep evals. And I wanted to give 695 00:26:49,919 --> 00:26:52,159 deep eval shout out. It's kind of the pi 696 00:26:52,159 --> 00:26:56,159 test of the LLM space. Um what we 697 00:26:56,159 --> 00:26:59,039 thoroughly enjoyed was the ability for 698 00:26:59,039 --> 00:27:01,600 us to create generic evaluations the GE 699 00:27:01,600 --> 00:27:04,159 valves where we would provide our metric 700 00:27:04,159 --> 00:27:07,440 names, our criteria against particular 701 00:27:07,440 --> 00:27:09,760 uh test cases. [snorts] 702 00:27:09,760 --> 00:27:12,320 So for example, this is just a 703 00:27:12,320 --> 00:27:14,880 nonsensitive example. It's a little bit 704 00:27:14,880 --> 00:27:17,279 uh of a joke. Migrate the staff 705 00:27:17,279 --> 00:27:19,760 cafeteria menu publisher to the cloud 706 00:27:19,760 --> 00:27:22,080 and providing some retrieval context 707 00:27:22,080 --> 00:27:23,919 into this agent saying here are our 708 00:27:23,919 --> 00:27:26,400 architectural principles and based off 709 00:27:26,400 --> 00:27:28,480 the previous agents response, how would 710 00:27:28,480 --> 00:27:31,440 we grade the faithfulness, relevancy, 711 00:27:31,440 --> 00:27:33,919 and whether or not we've reused our 712 00:27:33,919 --> 00:27:35,760 architecture before deciding to build a 713 00:27:35,760 --> 00:27:37,440 new component in our architecture. And 714 00:27:37,440 --> 00:27:38,799 you can see here in this very simple 715 00:27:38,799 --> 00:27:41,279 example, we've got 100% on all three. 716 00:27:41,279 --> 00:27:43,520 But I do want to share, consider 717 00:27:43,520 --> 00:27:45,039 carefully if your evaluation is giving 718 00:27:45,039 --> 00:27:47,279 you 100% on the board. You might have 719 00:27:47,279 --> 00:27:49,840 evaluation saturations. So perhaps think 720 00:27:49,840 --> 00:27:51,919 about whether your criteria needs to be 721 00:27:51,919 --> 00:27:54,240 more detailed. [snorts] So if you're 722 00:27:54,240 --> 00:27:56,559 getting started with skills, evaluate, 723 00:27:56,559 --> 00:27:58,480 check your artifact. 724 00:27:58,480 --> 00:27:59,919 Don't necessarily be happy with the 725 00:27:59,919 --> 00:28:02,640 output of your artifact being correct. 726 00:28:02,640 --> 00:28:04,480 Check the trace that it's called the 727 00:28:04,480 --> 00:28:06,880 tools and systems you expect it to call. 728 00:28:06,880 --> 00:28:08,799 Don't run the evaluation only once. 729 00:28:08,799 --> 00:28:10,559 Please run it multiple times to check if 730 00:28:10,559 --> 00:28:13,760 there's any variance. And we've talked 731 00:28:13,760 --> 00:28:15,679 about DPV val as an example library. 732 00:28:15,679 --> 00:28:18,000 That's something you can run before you 733 00:28:18,000 --> 00:28:20,320 deploy to your users. If you're running 734 00:28:20,320 --> 00:28:22,480 again the context of an agent uh towards 735 00:28:22,480 --> 00:28:25,360 a user group within a product as you're 736 00:28:25,360 --> 00:28:27,600 observing traces on your remote trace 737 00:28:27,600 --> 00:28:30,720 logs, any failures of evaluations that 738 00:28:30,720 --> 00:28:32,880 occur remotely, you can put inside a 739 00:28:32,880 --> 00:28:34,720 case library or as a new fixture within 740 00:28:34,720 --> 00:28:37,919 your local lane. Now to finish up, 741 00:28:37,919 --> 00:28:39,440 prompt engineering and prompting is 742 00:28:39,440 --> 00:28:42,320 extremely powerful. Extremely powerful. 743 00:28:42,320 --> 00:28:44,399 It gets you started, but hopefully you 744 00:28:44,399 --> 00:28:46,559 can see and from your own experiences 745 00:28:46,559 --> 00:28:48,399 perhaps feel that skills help expand 746 00:28:48,399 --> 00:28:50,880 your agent. And evaluation, which I've 747 00:28:50,880 --> 00:28:52,240 just touched on very briefly here, 748 00:28:52,240 --> 00:28:54,960 enables you to progress with confidence 749 00:28:54,960 --> 00:28:57,200 and that competence that we're seeking. 750 00:28:57,200 --> 00:28:59,440 Now, I had a bit of a uh how would you 751 00:28:59,440 --> 00:29:01,360 say controversial title where I was 752 00:29:01,360 --> 00:29:03,919 suggesting to you all stop prompting, 753 00:29:03,919 --> 00:29:05,679 start composing. But what I want to 754 00:29:05,679 --> 00:29:07,840 share with you all to conclude is my 755 00:29:07,840 --> 00:29:10,640 message is to recognize when to stop 756 00:29:10,640 --> 00:29:12,159 prompting 757 00:29:12,159 --> 00:29:14,880 without a boundary without a constraint 758 00:29:14,880 --> 00:29:18,000 and when to start composing with 759 00:29:18,000 --> 00:29:21,360 abstractions and smaller units of work. 760 00:29:21,360 --> 00:29:24,645 Thank you. [applause]