How we can improve Inference and Memory of AI

Hello Everyone,

The inference cost of AI is the major problems from all the existing companies today and also the memory of AI which gets diluted after sometime.

Do you all think improving the memory persistent layers( using vector DB, semantic findings and Graph DB) all together and also using draft models for inference can solve the problem , not completely but in someway.

What you guys think?

Hey There :waving_hand:

This problem has been growing over the past couple of years actually and I’ve seen many ideas to try to achieve some kind of cost optimization. I think the mentioned ideas and others mitigate the problem but I think the cost problem is getting bigger and bigger everyday. Just one man’s opinion, but I believe in the not-so-distant future we might have cost optimization become a requirement to work as an AI developer. But on a more brighter note, I do think the methods of optimising cost are growing and expanding as well.

yes i believe that’s definitely true and if we research properly we can actually do something about or on it